
NVIDIA today announced optimizations across all its platforms to accelerate Meta Llama 3, the latest generation of the large language model (LLM).
The open model combined with NVIDIA accelerated computing equips developers, researchers and businesses to innovate responsibly across a wide variety of applications.
Trained on NVIDIA AI Meta engineers trained Llama 3 on computer clusters packing 24,576 NVIDIA H100 Tensor Core GPUs, linked with RoCE and NVIDIA Quantum-2 InfiniBand networks.
To further advance the state of the art in generative AI, Meta recently described plans to scale its infrastructure to 350,000 H100 GPUs.
Putting Llama 3 to Work Versions of Llama 3, accelerated on NVIDIA GPUs, are available today for use in the cloud, data center, edge and PC.
From a browser, developers can try Llama 3 at ai.nvidia.com. It's packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.
Businesses can fine-tune Llama 3 with their data using NVIDIA NeMo, an open-source framework for LLMs that's part of the secure, supported NVIDIA AI Enterprise platform. Custom models can be optimized for inference with NVIDIA TensorRT-LLM and deployed with NVIDIA Triton Inference Server.
Taking Llama 3 to Devices and PCs Llama 3 also runs on NVIDIA Jetson Orin for robotics and edge computing devices, creating interactive agents like those in the Jetson AI Lab.
What's more, NVIDIA RTX and GeForce RTX GPUs for workstations and PCs speed inference on Llama 3. These systems give developers a target of more than 100 million NVIDIA-accelerated systems worldwide.
Get Optimal Performance with Llama 3 Best practices in deploying an LLM for a chatbot involves a balance of low latency, good reading speed and optimal GPU use to reduce costs.
Such a service needs to deliver tokens - the rough equivalent of words to an LLM - at about twice a user's reading speed which is about 10 tokens/second.
Applying these metrics, a single NVIDIA H200 Tensor Core GPU generated about 3,000 tokens/second - enough to serve about 300 simultaneous users - in an initial test using the version of Llama 3 with 70 billion parameters.
That means a single NVIDIA HGX server with eight H200 GPUs could deliver 24,000 tokens/second, further optimizing costs by supporting more than 2,400 users at the same time.
For edge devices, the version of Llama 3 with eight billion parameters generated up to 40 tokens/second on Jetson AGX Orin and 15 tokens/second on Jetson Orin Nano.
Advancing Community Models An active open-source contributor, NVIDIA is committed to optimizing community software that helps users address their toughest challenges. Open-source models also promote AI transparency and let users broadly share work on AI safety and resilience.
Learn more about how NVIDIA's AI inference platform, including how NIM, TensorRT-LLM and Triton use state-of-the-art techniques such as low-rank adaptation to accelerate the latest LLMs.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
01/04/2026
January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION
Douyin Users Can Now Create And Share Videos With Stun...
26/02/2026
The agreement ensures Europe's satellite-based augmentation continues enhanc...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/02/2026
With more than four decades of experience in radio broadcasting and live sports production, Daryl Doss, owner of Doss Technical Services and a contract engineer...
26/02/2026
BCNEXXT has deployed live HLG-based HDR playout capabilities within its Vipe platform, enabling broadcasters to integrate High Dynamic Range into live productio...
26/02/2026
TAG Video Systems (Booth W2323) will unveil new capabilities across its IP-native Realtime Media Platform at NAB 2026. New releases include visual service healt...
26/02/2026
IBC today announced a new strategic partnership with EIT Culture & Creativity the institutional partnership for culture and creativity, supported by the Europ...
26/02/2026
Clear-Com kept the action on track at Red Bull Shay'iMoto, an adrenaline-fueled motorsport spinning event that transformed the streets of Durban, South Afr...
26/02/2026
Harmonic (NASDAQ: HLIT) today announced that Alcom, a leading telco operator in Finland, is powering its next-generation white-label headend video service with ...
26/02/2026
Big Blue Marble, a provider of broadcast-grade, cloud-native video solutions for broadcasters, service providers, and content owners, today announced that it ha...
26/02/2026
New approach enables video service providers to deliver multiple live feeds on the same screen with lower costs and improved device compatibility
Broadpeak, a ...
26/02/2026
GeForce NOW's anniversary celebration reaches a chilling crescendo as Capcom...
26/02/2026
Final quarter revenues increase 7% year-on-year, with accelerating momentum in the second half
Space Services grows revenues by 6% year-on-year and records hig...
25/02/2026
With the Olympic Flag officially handed over to the organisers of the next Winter Games and the baton passed from Milano Cortina 2026 to French Alps 2030, the I...
25/02/2026
From a studio overlooking the Dolomites to workflows routed through Milan and into Salford, the BBC delivered a lean and mean operation for its Winter Games c...
25/02/2026
Warner Bros. Discovery (WBD) Sports is managing a huge network of channels acros...
25/02/2026
From its base in the northern Italian town of Cortina, Warner Bros. Discovery (W...
25/02/2026
In addition to 16:9-to-9:16 intelligent cropping for live video, Inference autom...
25/02/2026
Longtime rivals Floyd Money Mayweather Jr. (50-0, 27 KOs) and Manny PacMan P...
25/02/2026
The WNBA's Portland Fire and NWSL's Portland Thorns announce a groundbre...
25/02/2026
Multi-angle coverage, on-demand access to ultra-high-resolution video are provided for replays and clips across multiple distribution channels
The NHL and Cosm...
25/02/2026
The implementation standardizes an integrated workflow connecting ultra-high-res...
25/02/2026
Targeting a younger audience, creator-led network's Access Granted series hi...
25/02/2026
Alpha, the project's systems integrator, assisted in the workflow transformation
Tipping off the second half of the 2025-26 home schedule against the Houst...
25/02/2026
OCVIBE, the 100-acre mixed-use development transforming the area surrounding Hon...
25/02/2026
It's never been easier to customize your Spotify listening experience. Last year, we introduced more control over the way your playlist sounds, giving Premi...
25/02/2026
Hip-hop thrives on constant reinvention, with bold voices and fearless experimentation continually pushing the genre's boundaries. Every era brings new lead...
25/02/2026
L3Harris technicians recently completed a major mirror refurbishment for the U.S...
25/02/2026
This new offering helps solve for the need to move beyond traditional audience d...
25/02/2026
Gold-standard Gracenote content metadata will power Samsung's LLM-enabled entertainment search discovery experiences and more
NEW YORK February 25, 202...
25/02/2026
Afrobeats Icon Tiwa Savage Joins Forces with Berklee to Empower African Talent In collaboration with Berklee Global, the Tiwa Savage Music Foundation will hos...
25/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...