Sony Pixel Power calrec Sony

What's the ROI? Getting the Most Out of LLM Inference

09/10/2024

Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.

But with opportunities often come challenges.

Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.

To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.

Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.

We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.

NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.

This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.

Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.

These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.

Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.

Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.

Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.

For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.

The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.

For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.

Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.

Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
LINK: https://blogs.nvidia.com/blog/llm-inference-roi/...
See more stories from nvidia

Most recent headlines

16/12/2025

Stranger Things' Playlist Takeover Challenges Fans to Decode Clues and Unlock Exclusive Volume 2' Content

Hawkins has landed on Spotify, just in time for Stranger Things Season 5, Volume...

16/12/2025

Spotify and NAVER Bring Music Integrations and Premium Benefits to Korea

Wherever you are, your favorite music and audio content should go seamlessly with you. That's why Spotify has partnered with NAVER Corp, Korea's leading...

16/12/2025

Around the World With Spotify Wrapped: 2025 Fan Destinations You Had to See

2025 Wrapped arrived bigger and bolder than ever. This year's experience is designed to be ultra personal and shareable, with new features like Wrapped Part...

16/12/2025

L3Harris Delivers Most Powerful Thrusters for NASA's Lunar Gateway

Three 12-kilowatt Advanced Electric Propulsion System thrusters, supplied by L3Harris Technologies, form the core of Gateway's propulsion system. Pictured i...

16/12/2025

Palantir and L3Harris: Reindustrializing Defense Through AI-Powered Production

The challenge facing America's defense industrial base is not just about speed - its about rebuilding the foundation that makes speed possible. Our nations ...

16/12/2025

NFL Boosts Broadcast, Streaming Viewing Share in November

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

16/12/2025

Spanish Broadcaster Updates Playout With Pebble Solutions

SEVILLE, Spain Canal Sur, the public broadcasting service for Andalusia, Spain, has completed a total technology refresh based on Pebble's resilient, softwa...

16/12/2025

Telescript International Buys Prompter Software Company

NEW YORK Teleprompting hardware provider Telescript International has acquired all software code and intellectual property previously owned by Telescript West. ...

16/12/2025

Ookla: T-Mobile Is Fastest Fixed Wireless Access Provider

As cable operators face increased competition from 5G fixed wireless access providers, a new report from Ookla Research finds that T-Mobile is the FWA speed lea...

16/12/2025

Apple TV App for Android Now Supports Google Cast

Apple has announced a major upgrade to the Apple TV app for device owners outside the Apple ecosystem with news that the Apple TV app for Android now supports G...

16/12/2025

PT Telkom Satelit Indonesia and Space42 Explore Partnership to Extend 5G Direct-to-Device Services

Space42 grows Direct-to-Device partner ecosystem through a Memorandum of Underst...

16/12/2025

VEON Announces Release Date for Full Year and Fourth Quarter 2025 Results of Both VEON and Kyivstar

16 Dec 2025 VEON Announces Release Date for Full Year and Fourth Quarter 2025 R...

16/12/2025

VEON's Kyivstar Invests in Renewable Energy in Ukraine with Acquisition of Solar Power Company

16 Dec 2025 VEON's Kyivstar Invests in Renewable Energy in Ukraine with Acq...

16/12/2025

Emma Appleton, Fares Fares, Frida Gustavsson and Jakob Oftebro Star in Swedish Thriller Series Bytet'

Back to All News Emma Appleton, Fares Fares, Frida Gustavsson and Jakob Oftebro...

16/12/2025

Docu-reality 'My Korean Boyfriend' Gets a Trailer and Premiere Date: January 1st. Is Real Life Really like a K-drama?

Back to All News Docu-reality My Korean Boyfriend Gets a Trailer and Premiere D...

16/12/2025

Czech TV Elevates Video Streaming with Harmonic

Harmonic's XOS Advanced Media Processor Improves Streaming Video Quality and Boosts Viewer Engagement SAN JOSE, Calif. - Dec. 16, 2025 - Harmonic (NASDAQ: ...

16/12/2025

RT Sport Manager of the Year Nominees 2025 Revealed

RT Sport Awards 2025 live on RT One and RT Player at 8:05pm on Saturday 20 December. On Saturday 20 December live on RT One and RT Player at the earlier t...

16/12/2025

Music legend Brian Kennedy revealed as the twelfth and final contestant for Dancing with the Stars 2026

Singer -songwriter Brian Kennedy has been announced as the final celebrity dance...

15/12/2025

Harlem Globetrotters Celebrate 100th Anniversary With New Brand Campaign From The Famous Group

Harlem Globetrotters Celebrate 100th Anniversary With New Brand Campaign From Th...

15/12/2025

2026 Sundance Film Festival Reveals 54 Titles Selected for Short Film Program Presented by Ketel One Vodka

Top L-R: La Tierra Del Valor (The Home of the Brave), Mangittatuarjuk (The Gnawe...

15/12/2025

L3Harris to Provide Assured Communications for US Air Force's Survivable Airborne Operations Center Program

L3Harris will leverage 15 years of experience supporting the E-4B Nightwatch and...

15/12/2025

Arkansas TV Ends PBS Affiliation Amid Funding Cuts

CONWAY, Ark. In a notable example of how the loss of federal funding is forcing public stations to make massive cuts and operational changes, the statewide pub...

15/12/2025

PMVG Acquires WBPA-LD for WQED Pittsburgh

BOULDER, Colo. Public Media Venture Group (PMVG), Venture Technologies Group (VTG), and WQED have completed a multipart agreement that they say will significant...

15/12/2025

SES and WPDI Win Changing Lives Award, Connecting Youth to Digital Education in South Sudan and Uganda

Cape Town, November 13, 2025 - SES and International artist and humanitarian, Fo...

15/12/2025

SES, Abra Group Launch Multi-Orbit Inflight Connectivity

Luxembourg, December 15, 2025 - SES, a leading space solutions company, and Abra Group launched fast and reliable multi-orbit inflight connectivity service on t...

15/12/2025

How Rivian s Design Puts Drivers First-And Why That Matters

How Rivian s Design Puts Drivers First-And Why That Matters Published on Dec 15, 2025 Categories: Business Solutions LinkedIn Corporate Communications Sha...

15/12/2025

Space42 and Cobham Satcom Redefine L-band Performance with New Portfolio of Thuraya-4 Terminals

Space42 and Cobham Satcom completed the full range of advanced terminals for the...

15/12/2025

VEON's Beeline Kazakhstan Delivers First Starlink Direct to Cell Call in Central Asia

15 Dec 2025 VEON's Beeline Kazakhstan Delivers First Starlink Direct to Cel...

15/12/2025

U&GOLD Reveals Top 10 Topical Christmas Cracker Jokes for 2025!

Andrew Mountbatten-Windsor finds himself the topic of year's cracker jokes Oasis, David Harbour, Celebrity Traitors and Angela Rayner all feature in this y...

15/12/2025

Comscore Expands Cross-Platform Campaign Measurement to Include Audio and Social

Comscore Expands Cross-Platform Campaign Measurement to Include Audio and Social New capabilities strengthen cross-platform campaign reporting suite; CCR rebran...

15/12/2025

NVIDIA Acquires Open-Source Workload Management Provider SchedMD

NVIDIA today announced it has acquired SchedMD - the leading developer of Slurm, an open-source workload management system for high-performance computing (HPC) ...

15/12/2025

RT.ie Achieves Major Milestone: One Billion Page Views in 2025

RT .ie has reached one billion page views this year and is on track to finish 2025 2% ahead of last year. Average time spent on the site is up 3% on 2024, with ...

15/12/2025

How to Fine-Tune an LLM on NVIDIA GPUs With Unsloth

Modern workflows showcase the endless possibilities of generative and agentic AI on PCs. Of many, some examples include tuning a chatbot to handle product-supp...

13/12/2025

YouTube TV to Launch Genre Packages

In a move that will help it offer more flexible and less costly programming options, YouTube TV has announced that it will be launching YouTube TV Plans with mo...

13/12/2025

Magna Systems Finishes UHD, IP-based OB Truck For Singapore Network

SINGAPORE Magna Systems has designed, built and completed what is believed to be the first full UHD and IP-based OB truck in Southeast Asia for a Singapore medi...

12/12/2025

SVG Summit 2025 Preview: Everything You Need to Know for Next Week's Big Show in NYC

SVG Summit 2025 Preview: Everything You Need to Know for Next Week's Big Sho...

12/12/2025

Hailey Gates and Alia Shawkat Welcome You to the Village of Atropia

Hailey Gates at the Atropia premiere (photo by George Pimentel / Shutterstock for Sundance Film Festival)...

12/12/2025

Spotify and ATP Tour Launch First Episode of New Video Series

Last month, Spotify announced a new collaboration with the ATP Tour, the global governing body of men's professional tennis, aimed at bringing the next gene...

12/12/2025

Arkansas TV Drops PBS Affiliation Amid Funding Cuts

CONWAY, Ark. In a notable example of how the elimination of Federal federal funding is forcing public stations to make massive cuts and changes in the way they...

12/12/2025

Wisycom and DPA Microphones Appoint Rene Moerch as Group...

Wisycom and DPA Microphones announce the appointment of Ren Moerch as Group Product Director, Wireless, a strategic leadership role that will guide the combine...

12/12/2025

SMPTE Releases Updated Engineering Report on Artificial I...

SMPTE , the home of media professionals, technologists, and engineers, in conjuncture with the European Broadcasting Union (EBU) and the Entertainment Technolog...

12/12/2025

Keepit and Ingram Micro form strategic relationship in Po...

Keepit, the vendor-independent, cloud-native data protection provider, today announced a strategic go-to-market relationship in Poland with Ingram Micro, a lead...

12/12/2025

Atomos Enhances FUJIFILM GFX ETERNA 55 with RAW Capabilit...

Atomos announced the immediate availability of a new firmware update for its Ninja TX GO and Ninja TX monitor-recorders, unlocking Open Gate 48P RAW recording w...

12/12/2025

Professional Wireless Systems Provides Comprehensive RF S...

Professional Wireless Systems (PWS) once again played a critical role in delivering flawless wireless coordination and support at the 2025 Latin Grammy Awards a...

12/12/2025

AIMS Announces Inaugural IPMX Product Testing and Certifi...

The Alliance for IP Media Solutions (AIMS), together with the Video Services Forum (VSF), the Advanced Media Workflow Association (AMWA) and the European Broadc...

12/12/2025

DHD Gears for Hamburg Open 2026 with Latest Audio Product...

DHD audio will demonstrate the latest additions to its range of digital audio production solutions on Booth 321 in Hall B6 at Hamburg Open 2026. The show will b...

12/12/2025

Chaos Brings macOS Support and AI Tools to V-Ray for Blen...

Chaos today announces the release of V-Ray for Blender, update 2, bringing its award-winning rendering technology to even more Blender users by adding support f...

12/12/2025

UltraLEDs Launches Precision LED Tape for Professional Fi...

Lighting specialist UltraLEDs has launched Precision LED Tape, a high-CRI lighting solution designed specifically for professional film, TV, and studio use. P...