Sony Pixel Power calrec Sony

What's the ROI? Getting the Most Out of LLM Inference

09/10/2024

Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.

But with opportunities often come challenges.

Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.

To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.

Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.

We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.

NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.

This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.

Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.

These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.

Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.

Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.

Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.

For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.

The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.

For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.

Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.

Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
LINK: https://blogs.nvidia.com/blog/llm-inference-roi/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

01/04/2026

DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION

January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION Douyin Users Can Now Create And Share Videos With Stun...

17/03/2026

FCC Announces TV Translator Call Sign Changes

Share Copy link Facebook X Linkedin Bluesky Email...

17/03/2026

2026 NAB Show Offering Free Show Floor Passes to Creators

Share Copy link Facebook X Linkedin Bluesky Email...

17/03/2026

QuickLink's Latest StudioEdge Models to Make North American Debut at NAB 202

QuickLink's Latest StudioEdge Models to Make North American Debut at NAB 202 Brie Clayton March 16, 2026 0 Comments The Multi-platform Remote Gues...

17/03/2026

Frankenstein Graded with DaVinci Resolve Studio

Frankenstein Graded with DaVinci Resolve Studio Brie Clayton March 16, 2026 0 Comments Sonnenfeld enhances the controlled interplay between warm and c...

17/03/2026

New Voyavox from Link Electronics with Real-Time Speech-to-Text Captioning to be Featured in NAB Booth #W2910

New Voyavox from Link Electronics with Real-Time Speech-to-Text Captioning to be...

17/03/2026

Berklee City Music Stewards META Fellowship Supporting Massachusetts Music Educators

Berklee City Music Stewards META Fellowship Supporting Massachusetts Music Educa...

17/03/2026

Snap Decisions: How Open Libraries for Accelerated Data Processing Boost A/B Testing for Snapchat

The features on social media apps like Snapchat evolve nearly as fast as what...

17/03/2026

GTC Spotlights NVIDIA RTX PCs and DGX Sparks Running Latest Open Models and AI Agents Locally

The paradigm of consumer computing has revolved around the concept of a personal...

16/03/2026

DAZN to Stream NCAA Men's and Women's Basketball Tourneys Free in Select International Markets

DAZN will allow fans in select international territories to watch the NCAA men&#...

16/03/2026

IDM and Skate Board Association Announce Arena and Training Complex Planned for Big Bear Lake

IDM and The Skate Board Association (SBA) have announced a partnership with Coop...

16/03/2026

NAB 2026: Solid State Logic Introduces ST 2110-to-Dante Converter

Solid State Logic (SSL) will debut the Net I/O ST 2110 Bridge at NAB 2026 (booth C6907), a standalone unit that converts between ST 2110 and Dante audio formats...

16/03/2026

NAB 2026: Marshall Electronics Launches First 4K All-IP Weatherproof NDI Camera

Marshall Electronics (Booth C8339) is introducing its first all-IP 4K POV camera, the CV574-WP, at NAB 2026. The camera carries an IP67 weatherproof rating for ...

16/03/2026

Sony Expands Camera Authenticity Solution to Support Video

Sony Electronics' Camera Verify (beta), a feature of its Camera Authenticity Solution which enables news organizations to share content authenticity informa...

16/03/2026

FloSports and Storied Sports Partner on Women's and College Sports Content

FloSports has announced a partnership with Storied Sports, a content and IP studio founded by former espnW and The Players' Tribune executives, to develop s...

16/03/2026

Montreux Jazz Festival Names Gravity Media as A/V Production Provider

Montreux Jazz Festival has announced a multi-year collaboration with Gravity Media, who will become the Festival's Audio Visual Production Provider followin...

16/03/2026

USSI Global Names Ralph Annunziata Senior Vice President of Operations

USSI Global, a provider of customized network, broadcast and digital signage systems and services, has announced Ralph Annunziata joined the company on Jan. 5 a...

16/03/2026

NAB 2026: Boland Communications to Show New OLED Displays and Video Wall Applications

Boland Communications (booth C3519) will exhibit at NAB Show 2026 in Las Vegas, ...

16/03/2026

ST 2110 On The Go? A Peek Inside BRISK, FOX Sports' Broadcast Remote IP Studio Kit

Built in partnership with Diversified, the system As the sports broadcast indus...

16/03/2026

Amagi Report: FAST Viewership Up 21%, AI Adoption Growing Across Media Operations

Global FAST (Free Ad-supported Streaming TV) viewership grew 21% year-over-year ...

16/03/2026

Behind The Mic: Netflix, NBC Tap Matt Vasgersian to Call MLB and Tony Dungy Is Out at NBC

Behind The Mic provides a roundup of recent news regarding on-air talent, includ...

16/03/2026

Cloudvocal launch the SonoFlex instrument mic

Promises studio-grade fidelity for the stage Cloudvocal have announced the launch of a new instrument mic designed for professional live performers and engi...

16/03/2026

Kenton reveal the USB Solo Mk2

Popular MIDI/CV converter & interface overhauled Kenton have announced the launch of the USB Solo Mk2, a new and improved version of their compact MIDI to C...

16/03/2026

Sonarworks Spring Sale

Running from 16-29 March 2026 Starting from today (16 March) and running until 29 March 2026, Sonarworks are offering discounts of up to 40% across their ra...

16/03/2026

L3Harris Carries Goddard's Legacy Into a New Era

Dr. Robert H. Goddard and a liquid oxygen-gasoline rocket in the frame from which it was fired on March 16, 1926, at Auburn, Massachusetts. Credit: NASA....

16/03/2026

L3Harris Military GPS Receiver Deliveries Surpass 100,000 Units

Precision-guided munitions shown in production illustrate one of many operational systems benefiting from modernized M-Code GPS, supporting assured positioning,...

16/03/2026

A+E Global Media Signs New Multiyear Deal With Nielsen Covering Audience Measurement and Media Intelligence

NEW YORK - March 16, 2026 - A E Global Media and Nielsen today announced a new,...

16/03/2026

aconnic ramping up delivery of commercial 100-gigabit system

aconnic AG (ISIN: DE000A0LBKW6), Munich, is delivering the first commercial 100-Gigabit systems following successful validation and certification for customer n...

16/03/2026

Spectrum Launches Multiview for March Madness

Share Copy link Facebook X Linkedin Bluesky Email...

16/03/2026

Ikegami To Spotlight Latest UNICAM 4K-UHD Cameras At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

16/03/2026

A+E Global Media Signs New Agreement With Nielsen

Share Copy link Facebook X Linkedin Bluesky Email...

16/03/2026

Shotoku Brings Broadcast-Grade Control to PTZ with New Au...

Shotoku USA, Shotoku Broadcast Systems' North American operation, will unveil significant additions to its platform at NAB 2026. Topping the list is the wor...

16/03/2026

Ikegami to Showcase Latest Generation TV Production Camer...

Ikegami USA will demonstrate the latest additions to its wide range of broadcast-quality cameras, controllers and monitors on Central Hall booth C3819 during th...

16/03/2026

[Updated] Carr Threatens Broadcast Licenses Over Iran War Coverage

Share Copy link Facebook X Linkedin Bluesky Email...

16/03/2026

ELEMENTS launches GRID at NAB Show 2026

ELEMENTS launches GRID at NAB Show 2026 Brie Clayton March 15, 2026 0 Comments North Hall, Booth N1717 ELEMENTS returns to NAB Show 2026, with an exp...

16/03/2026

Blackmagic Design Cameras Capture Artist Salavat Fidai's Micro Sculptures

Blackmagic Design Cameras Capture Artist Salavat Fidai's Micro Sculptures Brie Clayton March 15, 2026 0 Comments 6K sensor and open gate capabilit...

16/03/2026

DHD to Introduce Latest Generation Broadcast Audio Mixers at NAB 2026, Las Vegas

DHD to Introduce Latest Generation Broadcast Audio Mixers at NAB 2026, Las Vegas Brie Clayton March 15, 2026 0 Comments Hero image: Front of DHD RM1 P...

16/03/2026

VEON Files its 2025 Annual Report on Form 20-F

16 Mar 2026 VEON Files its 2025 Annual Report on Form 20-F Dubai and New York, March 16, 2026 - VEON Ltd. (Nasdaq: VEON), a global digital operator ( VEON'...

16/03/2026

Sky Commissions The 100 Day Split, A New Relationship Series Exploring What Time Apart Reveals About Lifelong Love

Six Couples. 100 Days Apart. One Question: Does Absence Make the Heart Grow Fond...

16/03/2026

Tina Fey, Jamie Dornan and Riz Ahmed announced as first three hosts of Saturday Night Live UK

Monday 16 March 2026 Tina Fey, Jamie Dornan and Riz Ahmed announced as first th...

16/03/2026

Netflix Has Released the Trailer for 'Love at Last,' Starring Eda Ece and Kaan Yildirim

Back to All News Netflix Has Released the Trailer for Love at Last, Starring Ed...

16/03/2026

All UK national newspapers move to private circulation reporting while remaining audited by ABC

The reporting option was introduced following extensive consultation with publis...

16/03/2026

Rose of Tralee Katelyn Cummins wins Dancing with the Stars 2026

After a nail-biting Grand Finale, Rose of Tralee Katelyn Cummins has been announced as the winner of Dancing with the Stars 2026. The four finalists each dance...

16/03/2026

New RT series Welcome to Moore Street gives a glimpse into life on iconic Dublin street

Welcome to Moore Street will begin on RT One and RT Player on Thursday 19 Marc...

15/03/2026

Visit ToolsOnAir at NAB Las Vegas 2026

Visit ToolsOnAir at NAB Las Vegas 2026 More Details:From April 19-22, join us at NAB Show Las Vegas in the North Hall, Booth N1258, for an exclusive preview of...