Sony Pixel Power calrec Sony

Wide Open: NVIDIA Accelerates Inference on Meta Llama 3

18/04/2024

NVIDIA today announced optimizations across all its platforms to accelerate Meta Llama 3, the latest generation of the large language model (LLM).

The open model combined with NVIDIA accelerated computing equips developers, researchers and businesses to innovate responsibly across a wide variety of applications.

Trained on NVIDIA AI Meta engineers trained Llama 3 on computer clusters packing 24,576 NVIDIA H100 Tensor Core GPUs, linked with RoCE and NVIDIA Quantum-2 InfiniBand networks.

To further advance the state of the art in generative AI, Meta recently described plans to scale its infrastructure to 350,000 H100 GPUs.

Putting Llama 3 to Work Versions of Llama 3, accelerated on NVIDIA GPUs, are available today for use in the cloud, data center, edge and PC.

From a browser, developers can try Llama 3 at ai.nvidia.com. It's packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Businesses can fine-tune Llama 3 with their data using NVIDIA NeMo, an open-source framework for LLMs that's part of the secure, supported NVIDIA AI Enterprise platform. Custom models can be optimized for inference with NVIDIA TensorRT-LLM and deployed with NVIDIA Triton Inference Server.

Taking Llama 3 to Devices and PCs Llama 3 also runs on NVIDIA Jetson Orin for robotics and edge computing devices, creating interactive agents like those in the Jetson AI Lab.

What's more, NVIDIA RTX and GeForce RTX GPUs for workstations and PCs speed inference on Llama 3. These systems give developers a target of more than 100 million NVIDIA-accelerated systems worldwide.

Get Optimal Performance with Llama 3 Best practices in deploying an LLM for a chatbot involves a balance of low latency, good reading speed and optimal GPU use to reduce costs.

Such a service needs to deliver tokens - the rough equivalent of words to an LLM - at about twice a user's reading speed which is about 10 tokens/second.

Applying these metrics, a single NVIDIA H200 Tensor Core GPU generated about 3,000 tokens/second - enough to serve about 300 simultaneous users - in an initial test using the version of Llama 3 with 70 billion parameters.

That means a single NVIDIA HGX server with eight H200 GPUs could deliver 24,000 tokens/second, further optimizing costs by supporting more than 2,400 users at the same time.

For edge devices, the version of Llama 3 with eight billion parameters generated up to 40 tokens/second on Jetson AGX Orin and 15 tokens/second on Jetson Orin Nano.

Advancing Community Models An active open-source contributor, NVIDIA is committed to optimizing community software that helps users address their toughest challenges. Open-source models also promote AI transparency and let users broadly share work on AI safety and resilience.

Learn more about how NVIDIA's AI inference platform, including how NIM, TensorRT-LLM and Triton use state-of-the-art techniques such as low-rank adaptation to accelerate the latest LLMs.
LINK: https://blogs.nvidia.com/blog/meta-llama3-inference-acceleration/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

18/06/2026

Hollywood Filmmakers Launch AI Production Platform Cascade

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

ATVA Blasts Deltavision Media for Demanding 'Egregious' Retrans Fees

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

MediaKind Completes Merger with Harmonic's Video Business

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

AJA Unveils Io Xpand Thunderbolt 5 Expansion Chassis At InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

Linda May Han Oh Receives Guggenheim Fellowship

Linda May Han Oh Receives Guggenheim Fellowship The Berklee professor, bassist, and composer will use the fellowship to debut Dreams of Knowing, an interdisci...

18/06/2026

Sync and Stream: GeForce NOW Connects to Members' Game Libraries Across Devices

Play favorite titles from popular game libraries, keep progress synced and jump ...

18/06/2026

At Cannes Lions, NVIDIA Partners Reshape Advertising and Marketing With AI

The digital era gave the advertising and marketing industry speed; the AI era is giving it autonomous operations. For companies building next-generation techn...

17/06/2026

EVS Achieves EcoVadis Gold Medal, Ranking in Top 5% of Companies Globally

EVS has announced it has received the EcoVadis Gold Medal for sustainability performance, ranking among the top 5% of companies globally in the Technology/Mid-S...

17/06/2026

Chyron Releases Weather 2.4 with Updated DataFlow Module

Chyron has released Weather 2.4, an update to its weather suite for broadcasters and meteorologists. The release focuses on enhancements to the DataFlow module,...

17/06/2026

VSiN Launches Best Bets TV, a 24/7 FAST Channel

VSiN, The Sports Betting Network, has announced the launch of Best Bets TV, a free ad-supported streaming TV (FAST) channel. The 24/7 channel is currently avail...

17/06/2026

Snapchat and Team Whistle Launch World Cup Creator Program Across Miami, New York, and Los Angeles

DAZN's Team Whistle and Snap Inc. have announced a creator program centered ...

17/06/2026

LiveU Supporting Broadcast and Public Safety Operations for Summer of Soccer

LiveU is providing video transmission technology for broadcasters, production companies, and public safety agencies across North America's busy Summer of So...

17/06/2026

NATAS Announces Two New Board Members and COO Promotion

The National Academy of Television Arts and Sciences (NATAS) has announced that Laurens Grant and Jacob Ullman have joined its Board of Directors. Chief of Staf...

17/06/2026

Akta Brings AI-First Cloud Video Workflows to Oracle Cloud Infrastructure

Akta, the AI-First SaaS video platform for modern broadcast and streaming operations, today announced that its video platform is now generally available on Orac...

17/06/2026

SVG Students to Watch: Nicholas Stafford, Texas Southern University

This recent graduate from Houston found inspiration in technical directing and now eyes a future career in sports production...

17/06/2026

InfoComm 2026: Audio-Technica Brings New Wireless, Networked Audio, and Certified Conferencing Solutions

Audio-Technica (booth C7959) arrives at InfoComm 2026 in Las Vegas with a slate ...

17/06/2026

SNS Positions AI Suite as On-Premise Answer to Video Indexing and Search Demands

SNS has published a guide addressing growing demand for AI-powered video indexing, transcription, facial recognition, and searchable metadata across media libra...

17/06/2026

Providius Announces Providius Direct for Network Troubleshooting on NETGEAR Pro AV Switches

Providius has announced Providius Direct, a workflow for investigating network i...

17/06/2026

NEP Launches NEP Platform Software Orchestration System, Selects Bridge Technologies VB440

NEP Group has announced the commercial availability of NEP Platform, a software ...

17/06/2026

BBright White Paper: SRT vs. RIST for Professional Video Contribution and Distribution

A new white paper examining Secure Reliable Transport (SRT) and Reliable Interne...

17/06/2026

Bango Research: 51% of Gen Z Say Highlights Are Replacing Live Games

New research from subscription bundling platform Bango finds that younger sports fans are increasingly consuming sport through highlights, clips, and social med...

17/06/2026

SMPTE Opens Entire Standards Library to Public at No Cost

SMPTE has announced that its complete Standards catalog is now freely available to the global media technology community, including all published SMPTE Standard...

17/06/2026

MediaKind Completes Merger with Harmonic's Video Business, Creating Independent Video Infrastructure Heavyweight

Harmonic has completed the sale of its Video Business to MediaKind for $145 mill...

17/06/2026

Omaha Productions to Produce 2026 World Series of Poker Main Event Coverage for ESPN

Omaha Productions will produce the 2026 World Series of Poker (WSOP) in Las Vega...

17/06/2026

SoundBridge 3.1.0 now available

New features, changes & bug fixes SoundBridge have just released another update for their remote collaboration-focused DAW - reviewed here in SOS March 2026...

17/06/2026

Fryette launch the Valvulator Mini

Valve-based front end for digital & modelling rigs The latest addition to Fryette's product range delivers a packed-down, pedalboard-friendly version of...

17/06/2026

The Biggest UK Pro Audio Show In 20 Years!

GearExpo UK - 27 June 2026 Sound On Sound are proud to announce GearExpo UK, a major new recording and music technology exhibition in London! This is the bi...

17/06/2026

Genelec introduce the 9402A System Management Device

SAM monitoring line-up gains Dante and AES67 support The latest expansion of Genelec's UNIO monitoring ecosystem introduces a new device that provides D...

17/06/2026

The R&SPR300 portable receiver from Rohde & Schwarz sets new standards in spectrum monitoring

The R&S PR300 portable receiver from Rohde & Schwarz sets new standards in spect...

17/06/2026

Elt Group and Rohde & Schwarz sign a cooperation agreement to explore commercial opportunities in electronic warfare and defense

Elt Group and Rohde & Schwarz sign a cooperation agreement to explore commercial...

17/06/2026

NABLF Graduates Its 2026 Broadcasting Leadership Training Class

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

ABC and ESPN Score Most-Watched NBA Finals Since 1998

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

Smartly, Roku Bring Social Performance Tools to CTV

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

Pliant Technologies Debuts Crewcom Flex at InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

The Immersive Supervisor Emerges as Hollywood's Next Production Role

The Immersive Supervisor Emerges as Hollywood's Next Production Role Brie Clayton June 17, 2026 0 Comments Above image: On set live immersive revi...

17/06/2026

Vertical Musical Playback Shot with Blackmagic PYXIS 6K

Vertical Musical Playback Shot with Blackmagic PYXIS 6K Brie Clayton June 17, 2026 0 Comments Large format sensor and DaVinci Resolve workflow used fo...

17/06/2026

DAZ 3D Launches New Game-Ready Character Assets Built for Modern Engines and Production Workflows

DAZ 3D Launches New Game-Ready Character Assets Built for Modern Engines and Pro...

17/06/2026

Spectrum Awards $1.1 Million in Digital Education Grants

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

XR Sports Alliance Adds New Members

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

AIMS Launches Free Online IPMX Training Series

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

Kiloview Partners with SFM to Expand AV-over-IP Solutions...

Montr al, Quebec, June 11, 2026 Kiloview, a leading provider of AV-over-IP and NDI -based video transmission solutions, today announced a distribution partner...

17/06/2026

Kiloview Launches U4 IP Video Dock Bringing Professional...

Changsha, China, June 15, 2026 Kiloview officially announced the launch of U4 IP Video Dock, a compact IP video decoder and output dock designed to bring prof...

17/06/2026

France Advances Europe's AI Future With NVIDIA Technologies

A year ago at NVIDIA GTC Paris at VivaTech, France laid out plans to advance local AI - from new AI factories and national compute capacity to open frontier mod...

17/06/2026

Half Year, Half Off! Ivory II Summer Sale is Here

Save 50% on All Ivory II Pianos and Collections Were halfway through the year, and we're cutting prices in half for the Ivory II Summer Sale! For a limite...

17/06/2026

Two in three fans will connect to venue WiFi this World Cup, Sky Business research reveals

Wednesday 17 June 2026 Two in three fans will connect to venue WiFi this World ...

17/06/2026

Visibility builds credibility - the tools you use every day...

Visibility builds credibility - the tools you use every day, now visible on your LinkedIn profile Published on Jun 17, 2026 Categories: Company News, Product ...