Sony Pixel Power calrec Sony

What's the ROI? Getting the Most Out of LLM Inference

09/10/2024

Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.

But with opportunities often come challenges.

Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.

To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.

Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.

We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.

NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.

This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.

Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.

These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.

Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.

Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.

Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.

For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.

The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.

For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.

Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.

Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
LINK: https://blogs.nvidia.com/blog/llm-inference-roi/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

07/10/2026

Dalet Flex LTS Delivers Smarter Media Operations from Ingest to Distribution

Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

25/08/2026

Open, Sustainable, Predictable: What Utah Scientific Is Bringing to IBC 2026

Open, Sustainable, Predictable: What Utah Scientific Is Bringing to IBC 2026 As broadcast leaders prepare to gather at the RAI Amsterdam for IBC 2026, the main ...

25/08/2026

The Big Shift: Inside NC State's Phased Move From SDI to SMPTE ST 2110

NC State's Matthew Byrd breaks down how an IP-based infrastructure is delivering greater flexibility and scalability for Wolfpack productions...

25/08/2026

Monumental Sports Network Adds DAZN as Distribution Partner for Monumental+ in the DMV

Deal expands the DTC reach in-market of Capitals, Wizards, Mystics games and oth...

25/08/2026

Jacksonville Jaguars Carlos Caceres on Producing Shows at an Ever-Evolving EverBank Stadium

The two-year renovation period will create a stadium of the future for the NFL f...

25/08/2026

Behind The Mic: Genie Bouchard Joins ESPN for 2026 US Open; NBC Sports Names College Football Announce Hosts

Behind The Mic provides a roundup of recent news regarding on-air talent, includ...

25/08/2026

LD Systems Installs L-Acoustics A Series at Shell Energy Stadium in Houston

Shell Energy Stadium in Houston has installed a new L-Acoustics A Series loudspeaker system, designed and installed by LD Systems, a Clair Global brand. The 20,...

25/08/2026

FloCollege Expands Distribution and Adds Features for 2026-27 Season

FloCollege is expanding its distribution and product offerings for the 2026-27 season, covering more than 20 NCAA partner conferences across Division I, II, and...

25/08/2026

TNA Wrestling and REVOLT Announce Content Partnership

TNA Wrestling and REVOLT have announced a content partnership bringing exclusive TNA Xplosion matches and content from TNA's library to REVOLT. Our indust...

25/08/2026

Gracenote: FAST Sports Programming Growth Outpaces Channel Growth

Gracenote, Nielsen's content intelligence business, has released its Q3 2026 Data Hub analysis showing that FAST sports programming is growing at more than ...

25/08/2026

Amagi Appoints Thomas dHrouville Head of Sales for France, Benelux, and Nordics

Amagi has appointed Thomas d'H rouville as Head of Sales for France, Benelux, and Nordics. Based in Paris, he will lead Amagi's growth strategy across t...

25/08/2026

IPC Appoints CCATech as Channel Partner in Brazil

IPC has appointed CCATech as its channel partner in Brazil, expanding the company's presence in Latin America. CCATech was founded in 2023 by Carlos Abrah o...

25/08/2026

FanConnect Introduces FC4 Multi-Output Media Player

FanConnect has announced the FC4, a multi-output media player designed to drive mid-sized video walls, multi-screen concession stands, and multi-display digital...

25/08/2026

JWX Launches JWX Content Hub for Publisher Video Distribution and Monetization

JWX has launched JWX Content Hub, a platform for publisher video that combines content transformation, social distribution, audience engagement, and monetizatio...

25/08/2026

Techex Appoints Stuart Almond as Chief Revenue Officer

Techex has appointed Stuart Almond as Chief Revenue Officer. Almond brings more than 25 years of experience across technology, media, and telecommunications, ha...

25/08/2026

IBC 2026: Grass Valley to Showcase Unified SDI, IP, and Software Infrastructure

Grass Valley will demonstrate infrastructure connecting SDI, SMPTE ST 2110, MXL, and software-based processing at IBC2026 (Booth 9.A01, RAI Amsterdam, September...

25/08/2026

Taking Control: How Professional Sports Teams Are Building Their In-Market Future as the RSN Model Fractures

Franchises are taking greater ownership of production, distribution, and the fan...

25/08/2026

The PRO Tour Takes Golf Broadcasts Wireless With BMG's REMI-Driven Production Model

The first-year circuit for former professional athletes leans on LiveU, Starlink...

25/08/2026

UNIVERSAL FILM AND FOCUS FEATURES RENEW EXCLUSIVE PARTNERSHIP WITH SUNDANCE INSTITUTE TO SUPPORT RISING FILMMAKERS

Announcing the 2026 Sundance Institute Breakthrough | Focus Features Fellows L...

25/08/2026

Soundgas third John Paul Jones auction now live

Auction Part 3 running until 9 September 2026 Soundgas have announced that the latest stage of their John Paul Jones Gear Auction is now live, and that bids...

25/08/2026

MONO Music Conference programme announced

Over 60 award-winning artists, songwriters, producers, engineers The MONO Music Conference (formerly Music Expo SF), San Francisco's premier event for m...

25/08/2026

Groove Synthesis update the 3rd Wave

All three versions receive firmware upgrade Groove Synthesis have just launched a free firmware update that kits the 3rd Wave out with two major new feature...

25/08/2026

Clear-Com FreeSpeak Cell Field Trial Demonstrates Cellular Intercom Capabilities...

eds3_5_jq(document).ready(function($) { $(#eds_sliderM519).chameleonSlider_2_1({...

25/08/2026

BMG Taps LiveU For REMI Coverage of Golf's Pro Tour

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Rate of Subscription Fee Hikes Drops for Netflix, Disney+ and Amazon

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Law&Crime Reconvenes Court TV on Amagi CLOUDPORT

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Akta Tech To Showcase AI-First Video Platform At IBC 2026

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Major Live Sports Events Bolster Broadcast Viewing in Q2 2026

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Lingopal debuts real-time AI translation advances for mul...

Lingopal, the AI-powered live translation and localization specialist, is unveiling new platform capabilities at IBC2026 (11 14 September, RAI Amsterdam), helpi...

25/08/2026

TV5MONDE combines AI and editorial expertise with Bminty...

Bminty is supporting TV5MONDE in the evolution of its editorial processes through a new AI-assisted press information generation solution. Combining the Louise ...

25/08/2026

Two Cultures. Millions Reached. One Berklee Connection.

Two Cultures. Millions Reached. One Berklee Connection. Meet Annie Dickinson and Zhengxi Zhang-the couple who turned their creative partnership into music car...

25/08/2026

Law and Crime Moves Court TV to Amagi CLOUDPORT for Unifi...

Amagi, the media industry cloud platform for unified broadcast, streaming, and monetization, today announced that Law&Crime has migrated its Court TV channel, s...

25/08/2026

Broadcast Solutions demonstrates world-class technology a...

Broadcast Solutions, a leading systems integrator and provider of innovative solutions for the broadcast and media industry, is showcasing the breadth of its ca...

25/08/2026

Leader boosts Test and Measurement capabilities for Denma...

Thatcham, UK 25 August 2026: Test & measurement innovator, Leader Electronics, has announced that leading broadcast rental company Minitech has purchased a Le...

25/08/2026

Appear and TVTEL power centralised VAR signal transport i...

Appear's X20 Platform enables visually lossless video contribution with sub-frame latency and full path redundancy for TVTEL's centralised VAR operation...

25/08/2026

Big Blue Marble to showcase next-generation 5G Broadcast...

World-first demonstration combines 5G Broadcast and Media over QUIC for seamless, ultra-low-latency hybrid media delivery Big Blue Marble will showcase new dev...

25/08/2026

MNC Software Unveils Unified Software Suite at IBC2026 to...

MNC Software, a global leader in network management and operational support systems tailored to the broadcast and media industry, will showcase its complete sof...

25/08/2026

VAB: Ads in Local News Boost Sales and Brand Perception

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

ESPN to Bump Up Streaming Subscription Prices

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

HPA Opens Call for Award Entries

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

CNN Returns to Air After Technical Difficulties

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Calrec to Showcase DMF Capabilities at IBC 2026

Share Copy link Facebook X Linkedin Bluesky Email...

25/08/2026

Imagine Communications to Showcase AI and MCP Strategy at...

Company to Demonstrate Practical AI Applications That Add Value Today and Longer-Term Path Toward Agentic AI At IBC2026 (Sept. 11 14, RAI Amsterdam, Stand #1.B...

25/08/2026

KRK Kreate Series Standardizes Monitoring at Miamis WDNA...

Serving South Florida with a unique blend of jazz, Latin jazz, salsa, reggae, blues, and culturally reflective programming, WDNA 88.9FM operates as both a tradi...

25/08/2026

GlobalM Brings GMX Broadcast Platform to Oracle Cloud Inf...

GlobalM announced that it is integrating its GMX software-defined media platform with Oracle Cloud Infrastructure (OCI), bringing together Oracle's global c...

25/08/2026

Anthony Boyle leads cast of thrilling new Sky Original Charmer

Siena Kelly, Connor Swindells and Sophie Wilde join Boyle in the eight-part series from Working Title Television and Story Films, directed by Emmy and BAFTA nom...

25/08/2026

Ella Purnell makes a killer comeback in Sweetpea Season 2: Teaser Trailer Revealed for Sky and STARZ Original

The critically acclaimed thriller will return to Sky in NovemberTuesday 25 Augus...

25/08/2026

Scottish Rugby Expands Managed Coach Communications Partnership With Riedel

Wuppertal August 25, 2026 Scottish Rugby Expands Managed Coach Communications Partnership With RiedelRiedel Communications today announced the continued expan...