
Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.
But with opportunities often come challenges.
Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.
To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.
Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.
We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.
NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.
This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.
Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.
These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.
Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.
Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.
Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.
For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.
The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.
For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.
Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.
Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
09/10/2026
September 10 2026, 06:00 (PDT) Dolby Expands Dolby OptiView Platform with New Capabilities at IBC 2026
New Sports Intelligence helps providers better unders...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
18/09/2026
DHD reports a high level of interest in its range of audio production equipment on all four days of the September 11-14 International Broadcasting Convention in...
18/09/2026
Sonnet Announces Thunderbolt 5 eGPU With 850-Watt Power Supply and Integrated Do...
18/09/2026
DeckBridge adds live Resolve playback and Fairlight controls to Stream Deck
Brie Clayton September 17, 2026
0 Comments
DeckBridge starting profiles ac...
18/09/2026
Learn to arrange tracks using Color With Live's Free Stencil Pack
Brie Clayton September 17, 2026
0 Comments
Stencils are Color With Live's or...
18/09/2026
Berklee Establishes Susan Tedeschi Scholarship in Honor of Grammy-Winning Alumna The scholarship is announced as Tedeschi returns to headline the Berklee Bloc...
18/09/2026
The inside scoop on attention-grabbing market assets from those in the know 17 September 2026
(L-R): Sophie Green, Todd Brown and Edwina Waddy
Find this episo...
17/09/2026
Lions vs. Bills will also be the first regular-season game at new Highmark Stadi...
17/09/2026
Teradek has announced Bolt 7, a wireless video and data system that adds an integrated 900MHz radio to the Bolt's zero-delay video pipeline. The 900MHz chan...
17/09/2026
DAZN has announced three appointments to its executive leadership team, all reporting to CEO Shay Segev effective immediately.
Patrick Delany has been appointe...
17/09/2026
Monumental Sports and Entertainment (MSE) has announced a leadership restructuring. Zach Leonsis has been named President and Vice Chairman, a newly created pos...
17/09/2026
CP Communications has been named Best of Florida for Audio-Visual Services, and its subsidiary Red House Streaming has been named Best Movie and Recording Studi...
17/09/2026
Scripps Sports and ION have announced a broadcast partnership for the Shriners Children's East-West Bowl, the nation's oldest college all-star football ...
17/09/2026
The 17th annual espnW: Women + Sports Summit, presented by Toyota, will take place October 14-16 at the Ojai Valley Inn in Ojai, California, with virtual attend...
17/09/2026
The Professional Audio Manufacturers Alliance (PAMA) and Shure Incorporated have...
17/09/2026
Lithuanian broadcaster LNK Group has completed a migration of playout and media management for its five national television networks - LNK, BTV, TV1, InfoTV, an...
17/09/2026
Net Insight has announced a pan-Asian live media network created with its regional partners, giving broadcasters, production companies, and media service provid...
17/09/2026
After a high school football injury, the Orlando native found a new path in video production, gaining experience on ESPN college football broadcasts, Tampa Bay ...
17/09/2026
Popular DAW gains free pitch-correction tool
Image-Line have teamed up with Antares to kit their popular DAW software out with a built-in pitch-correction p...
17/09/2026
Promises unprecedented realism, dynamic nuance and clarity
Gibson have teamed up with acoustic pickup experts Baggs, creating a next-generation' pickup...
17/09/2026
Two new libraries join percussion line-up
VSL (Vienna Symphonic Library) have just launched two new libraries that bring some intriguing new instruments int...
17/09/2026
Hardware & plug-in versions gain new features
Arturia have just released a free update that brings some significant new features to their MiniFreak synthesi...
17/09/2026
Respected journalist Catalina Fl rez to succeed Anton Enus as SBS World News pre...
17/09/2026
The Department of Sport, Arts and Culture (DSAC), in partnership with the Nation...
17/09/2026
Calrec brings leading audio solutions and long-term business value to IBC 2026 We're looking forward to meeting up with you all on 11-14th September in Amst...
17/09/2026
Streaming Content Ratings (SCR) Data Shifts to Daily Delivery Cadence, Mirroring...
17/09/2026
August brought further stabilization in total TV viewing time. Poles spent an average of 3 hours and 31 minutes a day watching video content on TV glass-just 2 ...
17/09/2026
Post-sports seasonal shift redistributes viewership across Poland, while streami...
17/09/2026
Television Content Analytics Pte Ltd (TVC), a Singapore-based deep-tech company specialising in sports technology, artificial intelligence and live broadcast in...
17/09/2026
Rise WIB, the global advocacy group championing gender diversity and career progression across the media technology industry, today announced the shortlist for ...
17/09/2026
RALEIGH, N.C. - Capitol Broadcasting Company (CBC) and Hurricanes Holdings annou...
17/09/2026
Lithuanian broadcaster deploys Playout X and Framelight X across five networks, combining on-premises software-defined playout with cloud-based disaster recover...
17/09/2026
Introducing Adobe Photoshop Elements & Premiere Elements 2027
Brie Clayton September 17, 2026
0 Comments
New tools make it easier than ever to enhance y...
17/09/2026
LA JOLLA, CA-While the brain orchestrates metabolic processes in the body, it al...
17/09/2026
New series of Aistear an Amhr in shares extraordinary stories behind Ireland'...
17/09/2026
The Late Late Show is back!
Liam Neeson, Siobh n McSweeney, Caitr ona Balfe a...
17/09/2026
The Neighbours Effect: Gen Z Brings Back Big Hair, Double Denim and 80s soap sty...
17/09/2026
The bold, new legal drama starring Dominic West and Sienna Miller launches on Sky and NOW on 2nd OctThursday 17 September 2026
Official trailer released for Sk...
17/09/2026
Every week we read headlines about new threats from AI. From taking over jobs to disrupting critical infrastructure, spreading misinformation or even destroying...
17/09/2026
Arqiva selected by SANZAAR to deliver global distribution of 2026 international ...
17/09/2026
CULVER CITY, CALIFORNIA This evening at the 78th Primetime Emmy Awards, Apple TV...
17/09/2026
A new creature-catching adventure is ready to stream from the cloud this week. Pawprint Studio's Aniimo arrives on GeForce NOW at launch, inviting gamers to...
17/09/2026
RT Statement:
RT is today confirming that its position regarding participation in the Eurovision Song Contest remains unchanged. RT will not participate in...
16/09/2026
Grass Valley has been recognized with a 2026 Technology & Engineering Emmy Award...
16/09/2026
Sports-media leaders explore how cloud, AI, and virtualization are transforming ...
16/09/2026
With a full crew on site in Springfield, deep access to players, and a growing d...
16/09/2026
John Wilson attends The History of Concrete premiere during the 2026 Sundance Film Festival at The Yarrow Theatre on January 22, 2026, in Park City, Utah. (Ph...
16/09/2026
Powerful soft synth receives free update
Wavea have just launched a new and improved version of their feature-packed soft synth, kitting it out with over 20...