
NVIDIA today announced optimizations across all its platforms to accelerate Meta Llama 3, the latest generation of the large language model (LLM).
The open model combined with NVIDIA accelerated computing equips developers, researchers and businesses to innovate responsibly across a wide variety of applications.
Trained on NVIDIA AI Meta engineers trained Llama 3 on computer clusters packing 24,576 NVIDIA H100 Tensor Core GPUs, linked with RoCE and NVIDIA Quantum-2 InfiniBand networks.
To further advance the state of the art in generative AI, Meta recently described plans to scale its infrastructure to 350,000 H100 GPUs.
Putting Llama 3 to Work Versions of Llama 3, accelerated on NVIDIA GPUs, are available today for use in the cloud, data center, edge and PC.
From a browser, developers can try Llama 3 at ai.nvidia.com. It's packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.
Businesses can fine-tune Llama 3 with their data using NVIDIA NeMo, an open-source framework for LLMs that's part of the secure, supported NVIDIA AI Enterprise platform. Custom models can be optimized for inference with NVIDIA TensorRT-LLM and deployed with NVIDIA Triton Inference Server.
Taking Llama 3 to Devices and PCs Llama 3 also runs on NVIDIA Jetson Orin for robotics and edge computing devices, creating interactive agents like those in the Jetson AI Lab.
What's more, NVIDIA RTX and GeForce RTX GPUs for workstations and PCs speed inference on Llama 3. These systems give developers a target of more than 100 million NVIDIA-accelerated systems worldwide.
Get Optimal Performance with Llama 3 Best practices in deploying an LLM for a chatbot involves a balance of low latency, good reading speed and optimal GPU use to reduce costs.
Such a service needs to deliver tokens - the rough equivalent of words to an LLM - at about twice a user's reading speed which is about 10 tokens/second.
Applying these metrics, a single NVIDIA H200 Tensor Core GPU generated about 3,000 tokens/second - enough to serve about 300 simultaneous users - in an initial test using the version of Llama 3 with 70 billion parameters.
That means a single NVIDIA HGX server with eight H200 GPUs could deliver 24,000 tokens/second, further optimizing costs by supporting more than 2,400 users at the same time.
For edge devices, the version of Llama 3 with eight billion parameters generated up to 40 tokens/second on Jetson AGX Orin and 15 tokens/second on Jetson Orin Nano.
Advancing Community Models An active open-source contributor, NVIDIA is committed to optimizing community software that helps users address their toughest challenges. Open-source models also promote AI transparency and let users broadly share work on AI safety and resilience.
Learn more about how NVIDIA's AI inference platform, including how NIM, TensorRT-LLM and Triton use state-of-the-art techniques such as low-rank adaptation to accelerate the latest LLMs.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
10/04/2026
The Invisible OPEX Killer: Is Your Server Room Dragging You Down? In the broadcast world, we talk a lot about uptime. We talk about talent retention, latency...
10/04/2026
Imagine Communications will showcase its multiviewer portfolio at NAB Show 2026 (April 19-22, Booth N1328, Las Vegas Convention Center), including Prismon and t...
10/04/2026
Chyron has released PRIME VSAR 2.3, an update to its virtual set and augmented reality solution for broadcast. The release adds compatibility with Unreal Engine...
10/04/2026
Techex will exhibit at NAB Show 2026 (Booth W2267, April 19-23, Las Vegas Convention Center), demonstrating new tx darwin features including consumer multiview,...
10/04/2026
NDI will exhibit at NAB Show 2026, demonstrating its IP video ecosystem through live partner integrations, NDI 6.3 features, AI metadata workflows, and creator ...
10/04/2026
FOR-A has announced the acquisition of all shares of Tamu Radiance Corporation, a new company spun off from the Information Equipment Business of Tamura Corpora...
10/04/2026
InSync Technology will showcase new and updated video conversion products at NAB...
10/04/2026
TNT Sports and DAZN have announced a partnership to air monthly boxing events in the United States under the brand The Fight. The series will be promoted in p...
10/04/2026
Panasonic Projector and Display has announced the SQ3 Series of 4K LCD displays as part of its MEVIX professional display portfolio. All sizes will be available...
10/04/2026
Amagi has announced the addition of Agentic Media Operations to its Amagi NOW platform, integrating AI reasoning agents across its media supply chain workflows ...
10/04/2026
LTN has announced enhancements to its global IP video network targeting broadcasters transitioning from satellite distribution. The updates come ahead of US fed...
10/04/2026
Daktronics has installed new LED displays at Yankee Stadium, upgrading the main centerfield board, two flanking boards, and two ribbon displays spanning the 200...
10/04/2026
Harmonic has announced updates to its hybrid streaming solution, including Model Context Protocol (MCP) connectivity for AI applications, cloud-native deploymen...
10/04/2026
MultiDyne Video and Fiber Optic Systems will introduce two new fiber transport products at NAB Show 2026 (Booth C4425, April 19-22): the FiberSaver-10G waveleng...
10/04/2026
Telos Alliance and ip-studio will demonstrate STUDIO ZERO, a cloud-hosted virtual studio, at NAB Show 2026. First introduced at NAB Show 2023, STUDIO ZERO integ...
10/04/2026
d&b solutions, a London-based audio-visual, lighting, and media integration grou...
10/04/2026
ARRI and SmallHD have announced a new expansion license for ARRI's Hi-5 and Hi-5 SX hand units that displays lens data overlays on supported SmallHD monitor...
10/04/2026
Roku and the Banana Ball Championship League (BBCL) have announced an exclusive streaming partnership to bring five BBCL games to the Roku Sports Channel in 202...
10/04/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
10/04/2026
The Peabody Awards don't just recognize great storytelling, they spotlight t...
10/04/2026
After launching the Spotify Podcast Awards in Mexico last year, we brought the fan-voted celebration to Paris this week for its first edition in France. Hosted ...
10/04/2026
Powered and unpowered live PA ranges upgraded
Yamaha have just refreshed four of their hugely popular PA speaker ranges, delivering significant improvements...
10/04/2026
Underlying plug-in & VI technology now available to others
UJAM's latest announcement sees the company open up' Gorilla Engine, the development pla...
10/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/04/2026
NEWPORT BEACH, Calif., April 10, 2026 Bitcentral, a leading provider of professional media solutions for broadcast and digital video, will showcase its latest...
10/04/2026
Ikegami to Introduce Expanded Range of Broadcast Production Solutions at NAB 202...
10/04/2026
AJA Debuts SMPTE ST 2110 and openGear Solutions Ahead of NAB 2026
Brie Clayton April 10, 2026
0 Comments
New gear and updates address evolving hybrid ...
10/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/04/2026
Frequency, the engine behind the worlds leading streaming television channels, today launched its AI platform for Frequency Studio, powering the entire channel ...
10/04/2026
What can I watch on UKTV and stream on U this week?
This week on UKTV and the free streaming service U, viewers can watch a range of new and returning programm...
10/04/2026
J nger Audio Joins EBU ADM Integration Group as Founding Member to Help Advance...
10/04/2026
Five-part Sky Original drama airs nightly on Sky Mix and Sky Atlantic from 20 Ap...
09/04/2026
Staines-upon-Thames, UK, 09, April, 2026 - Yospace, the trusted leader in Dynam...
09/04/2026
just:play pro 2026 and just:live pro 2026 Sneak Preview News for NAB 2026
More Details:At NAB 2026, ToolsOnAir will showcase just:play pro 2026 and just:live p...
09/04/2026
just:in mac pro 2026 - The Next Level of Professional Recording on macOS at NAB ...
09/04/2026
Zixi will demonstrate IP-based live video workflow solutions at NAB Show 2026 (Booth W2057).
The industry is moving quickly toward IP-based distribution as br...
09/04/2026
Global women's elite sports revenues are expected to reach at least $3 billi...
09/04/2026
Monitor engineer Gavin Tempany mixed Kylie Minogue s Tension Tour on a Solid Sta...
09/04/2026
KOKUSAI DENKI Electric America will exhibit at NAB Show 2026 (Booth C5507), debu...
09/04/2026
With the 2025-26 NBA regular season concluded and the playoffs beginning next we...