Sony Pixel Power calrec Sony

Wide Open: NVIDIA Accelerates Inference on Meta Llama 3

18/04/2024

NVIDIA today announced optimizations across all its platforms to accelerate Meta Llama 3, the latest generation of the large language model (LLM).

The open model combined with NVIDIA accelerated computing equips developers, researchers and businesses to innovate responsibly across a wide variety of applications.

Trained on NVIDIA AI Meta engineers trained Llama 3 on computer clusters packing 24,576 NVIDIA H100 Tensor Core GPUs, linked with RoCE and NVIDIA Quantum-2 InfiniBand networks.

To further advance the state of the art in generative AI, Meta recently described plans to scale its infrastructure to 350,000 H100 GPUs.

Putting Llama 3 to Work Versions of Llama 3, accelerated on NVIDIA GPUs, are available today for use in the cloud, data center, edge and PC.

From a browser, developers can try Llama 3 at ai.nvidia.com. It's packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Businesses can fine-tune Llama 3 with their data using NVIDIA NeMo, an open-source framework for LLMs that's part of the secure, supported NVIDIA AI Enterprise platform. Custom models can be optimized for inference with NVIDIA TensorRT-LLM and deployed with NVIDIA Triton Inference Server.

Taking Llama 3 to Devices and PCs Llama 3 also runs on NVIDIA Jetson Orin for robotics and edge computing devices, creating interactive agents like those in the Jetson AI Lab.

What's more, NVIDIA RTX and GeForce RTX GPUs for workstations and PCs speed inference on Llama 3. These systems give developers a target of more than 100 million NVIDIA-accelerated systems worldwide.

Get Optimal Performance with Llama 3 Best practices in deploying an LLM for a chatbot involves a balance of low latency, good reading speed and optimal GPU use to reduce costs.

Such a service needs to deliver tokens - the rough equivalent of words to an LLM - at about twice a user's reading speed which is about 10 tokens/second.

Applying these metrics, a single NVIDIA H200 Tensor Core GPU generated about 3,000 tokens/second - enough to serve about 300 simultaneous users - in an initial test using the version of Llama 3 with 70 billion parameters.

That means a single NVIDIA HGX server with eight H200 GPUs could deliver 24,000 tokens/second, further optimizing costs by supporting more than 2,400 users at the same time.

For edge devices, the version of Llama 3 with eight billion parameters generated up to 40 tokens/second on Jetson AGX Orin and 15 tokens/second on Jetson Orin Nano.

Advancing Community Models An active open-source contributor, NVIDIA is committed to optimizing community software that helps users address their toughest challenges. Open-source models also promote AI transparency and let users broadly share work on AI safety and resilience.

Learn more about how NVIDIA's AI inference platform, including how NIM, TensorRT-LLM and Triton use state-of-the-art techniques such as low-rank adaptation to accelerate the latest LLMs.
LINK: https://blogs.nvidia.com/blog/meta-llama3-inference-acceleration/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

03/04/2026

TNT Sports and CBS Sports To Reunite Michigan's Iconic Fab Five' for Special NCAA Men's Final Four Altcast on truTV and HBO Max

Michigan's Fab Five will reunite for an alternate presentation of the Mich...

03/04/2026

NAB 2026: Avid To Showcase Content Core and AI Workflow Innovations

Avid will exhibit at NAB Show 2026 (April 18-22, Booth N2226, Las Vegas Convention Center), demonstrating its Content Core platform and new AI-driven workflow c...

03/04/2026

MRMC Appoints Nick Barthee as Chief Operating Officer

Mark Roberts Motion Control (MRMC) has announced the appointment of Nick Barthee as Chief Operating Officer. The announcement follows MRMC's transition fro...

03/04/2026

Elite Media Technologies Selects Interra Systems' BATON for File-Based QC

Interra Systems has announced that Elite Media Technologies has selected its BATON file-based QC solution for media workflows. Elite Media Technologies speciali...

03/04/2026

Moldtelecom Deploys Ateme Technologies Across Full Streaming Workflow

Ateme has announced that Moldtelecom has deployed Ateme technologies across its streaming workflow, covering encoding, delivery, operations, and analytics. Mol...

03/04/2026

NAB 2026: Grass Valley To Demonstrate Framelight X Content Management

Grass Valley will demonstrate Framelight X, its content management platform, at NAB Show 2026. The platform connects capture, ingest, editing, and publishing in...

03/04/2026

NAB 2026: Encompass Digital Media and Techex Launch Cloud-Based Master Control Service for Live Events

Encompass Digital Media and Techex have announced a cloud-native Master Control ...

03/04/2026

Peacock Debuts Its New Vertical-Video Experience for Live NBA Games on Monday

Live Vertical Video automatically track the action on the court via AI technology and delivers a fully optimized, 9 16 live feed for viewers...

03/04/2026

Illinois Creative Team Captures Men's Final Four Run With Trust, Timing, and a Few Water Guns

As the Illini make their first trip to college basketball's biggest stage si...

03/04/2026

Hook 'Em: University of Texas Athletics Produces Digital Content for a First-Person Experience

After last summer's Softball National Championship victory and last week'...

03/04/2026

Don't Be Lame: Arizona Men's Basketball Social Team Aims To Catch the Attention of Wildcat Fans

The University of Arizona's Men's Basketball team has only loss twice th...

03/04/2026

During Packed Weekend, Van Wagner Covers In-Venue Shows for Championships Across Indianapolis

Eight games across four tournaments will be played in three venues; accommodatio...

03/04/2026

Ottawa Senators and Bell Media Extend Regional Broadcast Rights Agreement

The Ottawa Senators and Bell Media have announced a long-term rights extension for regional Ottawa Senators games on TSN and RDS. TSN Radio 1200 remains the exc...

03/04/2026

ESPN's Women's Final Four Playbook: Bigger Compound, Bigger Trucks, Biggest Production Yet

Massive production in Phoenix running out of Flagship Mobile unit, Features 50+ ...

03/04/2026

Electro-Harmonix release EHX Classics Bundle

Iconic guitar pedals now available in plug-in form Guitar effects experts Electro-Harmonix have teamed up with MixWave to turn a collection of their most pr...

03/04/2026

FAC launch Bandit 2

New multi-band AUv3 plug-in announced Fred Anton Corvest (FAC) offer an extensive range of AUv3 plug-ins and iOS/iPadOS Apps, and their multiband effects pr...

03/04/2026

Pulsar-23: 1984 from SOMA Laboratory

Just 84 units to be released in the US Experimental synthesizer and sound-machine extraordinaires SOMA Laboratory have revealed an upcoming special-edition ...

03/04/2026

Iconic Instruments release Model 350 Tube Preamp

Emulates the input section of an Ampex 350 One of the latest arrivals to the Iconic Instruments range delivers a new tube preamp plug-in inspired by the cir...

03/04/2026

TelevisaUnivision and Nielsen Agree to New Media Intelligence Deal Covering Streaming, National and Local TV, and Radio

New York April 2, 2026 TelevisaUnivision, the world's leading Spanish-la...

03/04/2026

2026 NAB Show Exhibitor Insight: Techex

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

TelevisaUnivision Signs New Nielsen Media Intelligence Deal

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

The WNET Group, JIB Launch NHK World-Japan in New York

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

NAB Leadership Foundation Welcomes New Board Members

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

EverPass Media Expands Distribution Deal with Netflix

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

Versant Acquires AI-Data Platform StockStory

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

CVP Grows European Footprint with Strategic Expansion in...

CVP, one of Europe's leading suppliers of professional video and broadcast solutions, today announces the launch of its new German operation and the formati...

03/04/2026

MRMC Announces Appointment of Chief Operating Officer

Mark Roberts Motion Control (MRMC) today announces the appointment of Nick Barthee as Chief Operating Officer, strengthening its leadership as the company conti...

03/04/2026

Net Insight Introduces Programmable Trust Boundaries for...

Net Insight introduces programmable Trust Boundaries that make live media interconnection predictable as traffic moves between facilities, networks and cloud en...

03/04/2026

Winning in the new media economy: Avid showcases AI-powered, connected intelligence to unlock media value at NAB Show 2026

Winning in the new media economy: Avid showcases AI-powered, connected intellige...

03/04/2026

NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibility and Loudness at NAB 2026

NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibilit...

03/04/2026

NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances

NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances Brie Clayton April 2, 2026 0 Comments PlayBox Neo will showcase the latest i...

03/04/2026

For Taku Hirano, Everything Is Connected

For Taku Hirano, Everything Is Connected From touring and composition to teaching and instrument design, the in-demand percussionist sees it all as one body o...

03/04/2026

Berklee Honors Humberto Ramirez with Master of Latin Music Award

Berklee Honors Humberto Ramirez with Master of Latin Music Award The alumnus and acclaimed trumpeter is honored for his influence as a performer, composer, an...

03/04/2026

VIZ Media Lands Rumiko Takahashi's MAO, Sets April 4 Premiere on Hulu in the U.S. and Disney+ in Select International Markets

VIZ Media Lands Rumiko Takahashi's MAO, Sets April 4 Premiere on Hulu in the...

03/04/2026

Competition Heats Up with Intrigue and Spices: Netflix Unveils Trailer for New Drama Made with Love'

Back to All News Competition Heats Up with Intrigue and Spices: Netflix Unveils...

02/04/2026

HBO and NFL Films Announce Hard Knocks: Training Camp with the Seattle Seahawks, Debuting August 11

HBO and NFL Films have announced Hard Knocks: Training Camp with the Seattle Sea...

02/04/2026

NAB 2026: Haivision Unveils Makito ONE Video Transport Platform

Haivision has announced the Makito ONE, a single-blade video encoding and decoding platform, at NAB Show 2026. The platform combines dual-channel video encoding...

02/04/2026

NAB 2026: Telestream Introduces UP.Lens Cloud-Based Multiviewer and Monitoring Service

Telestream has introduced UP.Lens, a cloud-based multiviewer and monitoring serv...

02/04/2026

NAB 2026: MRMC to Showcase Robotic Camera Technology and Mark 60th Anniversary

Mark Roberts Motion Control (MRMC) will exhibit at NAB Show 2026 (Booth C5220, April 19-22, Las Vegas Convention Center), marking the company's 60th anniver...

02/04/2026

NAB 2026: Net Insight Introduces Programmable Trust Boundaries for Live Media Interconnection

Net Insight has introduced programmable Trust Boundaries, a feature integrated i...

02/04/2026

NAB 2026: Bitmovin Adds SGAI Support to Playback Products

Bitmovin has announced support for SGAI (Server-Guided Ad Insertion) in its playback products, using HLS interstitials. SGAI combines elements of client-side an...

02/04/2026

Binghamton University Athletics Adds Riedel SimplyLive RiMotion R12 for Student-Run Productions

Riedel Communications' SimplyLive RiMotion R12 replay system is supporting B...

02/04/2026

NAB 2026: LTN and Ateme Announce Integration of Video Processing with IP Transport

LTN, a managed IP video transport company, and Ateme, a video compression and de...

02/04/2026

TDF Expands Channel Capacity on Terrestrial Broadcast Network with Harmonic

Harmonic has announced that TDF, a broadcast infrastructure operator in France, has deployed Harmonic's XOS Advanced Media Processor and ProStream X Video S...

02/04/2026

United Rugby Championship Reports First-Year Results with Eluvio Streaming Platform

Eluvio and the United Rugby Championship (URC) have announced first-year results...