Sony Pixel Power calrec Sony

What's the ROI? Getting the Most Out of LLM Inference

09/10/2024

Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.

But with opportunities often come challenges.

Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.

To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.

Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.

We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.

NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.

This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.

Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.

These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.

Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.

Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.

Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.

For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.

The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.

For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.

Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.

Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
LINK: https://blogs.nvidia.com/blog/llm-inference-roi/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

08/04/2026

Synamedia unveils Senza Ignite to transform deployed devi...

Leading video software provider Synamedia today announced Senza Ignite, a cloud-based platform that transforms existing connected devices through a single firmw...

08/04/2026

Appear works with Corus to transform national news contri...

Modernised architecture increases operational control and agility, while reducing cost and simplifying nationwide media transport across affiliates and news bur...

08/04/2026

LynTec to Showcase High Density Power Control Solutions f...

Company to Feature DMX Dual Relay, High-Density 48-Channel Relay Panels, and New Private Label Power Program LynTec, a leading manufacturer of innovative ele...

08/04/2026

BCE integrates BCNEXXT Vipe to power cloud playout in Med...

Broadcasting Center Europe (BCE), a European media technology and services partner, today announces a technology partnership with BCNEXXT, a Netherlands-based c...

08/04/2026

Encompass Strengthens Global Media Operations with Expans...

Encompass Digital Media today announced the expansion of its Master Control Room (MCR) facility in Riga, Latvia, establishing it as one of the most advanced and...

08/04/2026

Synamedia lights up Las Vegas with innovations that will...

Leading video solutions provider Synamedia is back at The 2026 NAB Show with innovations that connect tomorrows audiences today. On display will be a range of t...

08/04/2026

COW Jobs: Director Needed for WWII Indie Docudrama

COW Jobs: Director Needed for WWII Indie Docudrama Brie Clayton April 8, 2026 0 Comments Director Needed for WWII Indie Docudrama March 17, 2026COW ...

08/04/2026

ESPN Expands Global Reach on Disney+

Share Copy link Facebook X Linkedin Bluesky Email...

08/04/2026

Broadcast Industry Initiatives Promise To Speed Adoption of ST 2110

Share Copy link Facebook X Linkedin Bluesky Email...

08/04/2026

Remembering Television's Role in the U.S. Bicentennial

Share Copy link Facebook X Linkedin Bluesky Email...

08/04/2026

Synamedia Unveils AI by Quortex

Share Copy link Facebook X Linkedin Bluesky Email...

08/04/2026

Carrie Healey Joins NAB as Vice President of Communications

Share Copy link Facebook X Linkedin Bluesky Email...

08/04/2026

Amagi Launches Newspulse

Share Copy link Facebook X Linkedin Bluesky Email...

08/04/2026

Marshall Electronics Unveils CV376 NDI HX3 HDMI POV Camer...

Marshall Electronics introduces its latest POV camera with the new CV376 Compact NDI |HX3 HDMI POV Camera at NAB 2026 (Booth C8339). Designed for versatile 4K c...

08/04/2026

Zixi Showcases Interoperable Live Video Workflows and Sat...

Zixi, a leader in live video delivery and workflow orchestration, enables interoperable live video workflows across multi-vendor environments, helping broadcast...

08/04/2026

Kathryn Thomas questions if we can be Young Forever in new RT documentary

Young Forever: The Death of Ageing? airs Monday 13 April and 20 April on RT One and RT Player Around the world, the race to beat ageing is on. From cutting ...

08/04/2026

April 07, 2026

Scripps Research scientists uncover new mechanism cancer cells use to survive DNA damage Discovery reshapes understanding of how tumor cells repair broken DNA, ...

07/04/2026

NAB 2026: Ikegami USA To Launch Two New Viewfinders, Refine UHD Cameras

Ikegami USA will launch a refinement to the UHK-X700 and UHK-X750 3-CMOS -in. UHD cameras in the UNICAM XE series plus two new 7-in. viewfinders at NAB 2026 in...

07/04/2026

NAB 2026: NAB Leadership Foundation to Host Technology Students Career Mixer

The NAB Leadership Foundation and NAB PILOT will host a career mixer for technology students at NAB Show 2026 on April 18, 5-6:30 p.m., North Hall, Room N225/22...

07/04/2026

Survey: AVoIP Adoption Accelerating, With Interoperability and Security as Key Drivers

Audinate Group Limited and Futuresource Consulting have published results from a...

07/04/2026

NAB 2026: Eluvio Announces Commercial Availability of Content Fabric Bucharest Release

Eluvio has announced the commercial availability of its Content Fabric Bucharest...

07/04/2026

NAB 2026: Eluvio Introduces Inline AI Video Intelligence and Updated EVIE

Eluvio has unveiled a new architecture for video AI and an updated Eluvio Video Intelligence Editor (EVIE) ahead of NAB Show 2026. Eluvio AI runs analysis and i...

07/04/2026

Professional Fighters League Partners With Sky New Zealand for Exclusive Broadcast Rights

The Professional Fighters League (PFL) has announced a deal with Sky New Zealand...

07/04/2026

NBA, Enjoy Basketball' To Produce Live Game Altcast, Enjoy the NBA Trivia Show as Part of New Multi-Platform Collab

The NBA and Enjoy Basketball, the digital media company co-founded by YouTube cr...

07/04/2026

Beyond Golf: PGA of America, Ko-Mar Productions Offer Clients a Customizable Production Studio

The 4,000-sq.-ft. space in Frisco, TX, has produced live and packaged programmin...

07/04/2026

A New POV: RefCam's Rise in the Bundesliga Signals a Potential New Era for Soccer Broadcasts

What began as a referee training tool is evolving into a powerful production ass...

07/04/2026

NAB 2026: Manifold Technologies to Join NEP Platform as Deployable Application

Manifold Technologies will announce at NAB Show 2026 (Booth C.1808) that its manifold CLOUD platform will be available as a deployable application within NEP Pl...

07/04/2026

Shaquille O'Neal and TNT Sports to Launch DUNKMAN Professional Dunk League in Summer 2026

Shaquille O'Neal, Authentic Brands Group, and TNT Sports, in partnership wit...

07/04/2026

NAB 2026: Telos Alliance Unveils Omnia XII Audio Processor

Telos Alliance will debut the Omnia XII, a new FM/HD/DAB audio processor, at NAB Show 2026 in Las Vegas. Built on a 2RU hardware platform, Omnia XII features a...

07/04/2026

NAB 2026: AWS to Showcase AI and Cloud Media Technologies

Amazon Web Services (AWS) will exhibit at NAB Show 2026 (April 18-22, Las Vegas Convention Center, Booth W1701), with demonstrations, speaking sessions, and int...

07/04/2026

NAB 2026: Synamedia to Demonstrate AI by Quortex

Synamedia will demonstrate AI by Quortex at NAB Show 2026, a framework that applies AI capabilities to video workflows on demand rather than continuously. The s...

07/04/2026

NAB 2026: TVNewsCheck Announces 2026 Women in Technology Award Honorees

TVNewsCheck will present its 15th annual Women in Technology Awards on Tuesday, April 21, at 5 p.m. PT at NAB Show, in the Media and Entertainment Theater (W146...

07/04/2026

NAB 2026: Akta to Showcase AI Video Platform

Akta will demonstrate its AI video platform at NAB Show 2026, highlighting new capabilities in media processing and vertical video formatting alongside its exis...

07/04/2026

ESPN Expands to Disney+ in Europe and Select Asia-Pacific Markets

ESPN and Disney have announced the launch of ESPN on Disney in Europe and select Asia-Pacific markets, bringing the offering to 53 countries and territories a...

07/04/2026

Announcing the 2026 Sundance Institute Native Lab Fellows

LOS ANGELES, CA, April 7, 2026 - The nonprofit Sundance Institute announced today the fellows selected for the 2026 Native Lab, the signature initiative of the ...

07/04/2026

Prompted Playlist Levels Up to Include Podcasts, Helping You Explore More Interests and Curiosities

Starting today, Prompted Playlist is expanding beyond music to now include podca...

07/04/2026

VSL release Synchron Harpsichord (Blanchet)

Launched alongside piano promotion The latest instalment in VSL's Synchron Series line-up captures the sound of a faithful copy of a Fran ois tienne Bl...

07/04/2026

Roland announce SPD-SX Pro Version 2.0

Flagship sampling pad gets an update Roland's flagship sampling pad has just received a major update that kits it out with an array of new features and ...

07/04/2026

Sound Devices unveil the Astral Mini Plus

Popular compact wireless system upgraded Sound Devices have recently introduced a new and improved version of their compact wireless transmitter bodypacks, ...

07/04/2026

YEP November 2025 Newsletter

The November 2025 YEP Newsletter highlights a recent YEP Coffee Chat, offering members the chance to connect with industry professionals in an informal setting ...

07/04/2026

YEP December 2025 Newsletter

The December 2025 YEP Newsletter includes a Spotlight on Emily Vail, showcasing her career journey and work in the industry, alongside Mentorship Reflections by...

07/04/2026

US Space Force Selects L3Harris to Strengthen America's Defense with Advanced Space Surveillance

Ground-Based Electro-Optical Deep Space Surveillance (GEODSS) telescope operated...

07/04/2026

Haivision Unveils Makito ONE Live Video Contribution Platform

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

Neutrik To Unveil TRUE1 Data Connector Series At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

Chris Welcker Deployed Full DPA Arsenal to Record Live Mu...

Catgut Sound Owner and Production Sound Mixer Chris Welcker, CAS, has built a career at the intersection of music and film. A former musician and composer, Welc...