
Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.
But with opportunities often come challenges.
Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.
To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.
Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.
We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.
NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.
This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.
Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.
These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.
Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.
Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.
Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.
For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.
The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.
For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.
Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.
Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
09/10/2026
September 10 2026, 06:00 (PDT) Dolby Expands Dolby OptiView Platform with New Capabilities at IBC 2026
New Sports Intelligence helps providers better unders...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
24/09/2026
Broadcast Management Group (BMG) is supporting Virginia Public Media (VPM), the ...
24/09/2026
RED Digital Cinema's RED V-RAPTOR [X] and RED V-RAPTOR XL [X] cameras with the Phantom Track feature are the first cameras to receive GhostFrame Pro Certifi...
24/09/2026
Bringing together the full family of networks, FOX Sports is also adding Tubi, F...
24/09/2026
Virginia Polytechnic Institute and State University (Virginia Tech) has unveiled...
24/09/2026
The blog collects SVG's reporting on AI across U.S. and European sports broadcasting
Since SVG launched the AI Innovation Lab in April to advance the devel...
24/09/2026
SVG's annual TranSPORT conference, which will be held on Tuesday, Oct. 20 in...
24/09/2026
The self-built Master Control Room has cut streaming operations incident-respo...
24/09/2026
Ortana Media Group, the software systems integrator behind the Cubix workflow platform, has connected AudioShake's audio separation models to ScorePlay thro...
24/09/2026
AudioShake has released Multi-Speaker 2.0, the company's newest Multi-Speaker model, which removes 75% more noise and offers 32% better separation of overla...
24/09/2026
By Jon Morgan, Founder and Executive Producer, Ryval Studios
For Major League B...
24/09/2026
Spectrum SportsNet has revealed its comprehensive programming and live game broa...
24/09/2026
The Charlotte Hornets has unveiled new broadcast partnerships with Gray Media an...
24/09/2026
DAZN will be the exclusive direct-to-consumer streaming partner for the Detroit Pistons beginning with the 2026-27 season. Subscriptions are available at DAZN.c...
24/09/2026
The National Hockey League (NHL) has made an agreement with Spectrum that will allow Spectrum TV customers in the Minnesota Wild local broadcast territory to wa...
24/09/2026
The Carolina Hurricanes has reached a multi-year distribution agreement to provi...
24/09/2026
The St. Louis Blues has introduced Blue Note as the name of the new, exclusive ...
24/09/2026
Blue Jackets Hockey Network, the television home for Blue Jackets hockey, is coming to Spectrum TV under a new carriage agreement between the National Hockey Le...
24/09/2026
The next generation of multitrack recording for live sound
The latest version of Waves' dedicated live multitrack recording software delivers an array ...
24/09/2026
An idea machine for musicians
Described as an idea machine for musicians , the latest addition to Polyend's growing collection of guitar pedals combin...
24/09/2026
Alignment tool gains multitrack layout
Synchro Arts have just launched the latest version of their renowned time-alignment plug-in, and it's said to be ...
24/09/2026
Synth gains DAW Control & Ctrl-e integration
Expressive E have just released a major firmware update that brings some interesting new features to the Osmose...
24/09/2026
Rohde & Schwarz expands its system amplifier portfolio with the R&S SAM200 for a...
24/09/2026
September 24, 2026
The Hitachi Startup Challenge FY26 invites innovative startu...
24/09/2026
Download Now!
To receive a short case study showcasing NHK Technologies, Inc. (NT) s audio setup, please enter the following details:
--
First Name* Fir...
24/09/2026
Manifold Technologies Appoints Ryhaan Williams as CCO for North America
Brie Clayton September 24, 2026
0 Comments
Brings three decades of broadcast e...
24/09/2026
UNIFACHA Expands Educational Program with DaVinci Resolve Studio
Brie Clayton September 24, 2026
0 Comments
Brazilian university empowers next generat...
24/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
24/09/2026
Viewers watching videos from the Austrian Parliament online will soon find it easier to verify the authenticity of the content. Big Blue Marble, an internationa...
24/09/2026
nsign, the digital signage SaaS platform built around its Simplify Complexity principle, is now officially certified for BRAVIA Professional Displays from Son...
24/09/2026
Encompass Digital Media today announced the successful migration of multiple Viaplay free-to-air channels, including TV3, TV6, TV8 and TV10, to its Altitude Sch...
24/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
24/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
24/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
24/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
24/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
24/09/2026
Music Supervisor, Composer, and Mixing Engineer Santiago Uribe has built a career around producing music for film and television, contributing to more than 40 p...
24/09/2026
The North American Broadcasters Association (NABA) will hold a full-day cybersecurity summit focused on defending against growing cyber threats facing the broad...
24/09/2026
London, 24th September 2026: UKTV has commissioned an eighth series of Canal Boat Diaries (8x60') for U (u.co.uk) and U&YESTERDAY, with waterways explorer R...
24/09/2026
September 24th, 2026 TRIBECA FESTIVAL AND AT&T OPEN SUBMISSIONS FOR 10TH ANNIVE...
24/09/2026
Reinaldo Marcus Green, director of six-time Academy Award-nominated King Richard...
24/09/2026
From the buzz of your phone when a message arrives to surgeons being able to feel' the patients they are examining from hundreds of kilometres away, these ...
24/09/2026
As Q4 approaches, retailers enter the most critical sales window of the year the golden quarter.
Stretching from October through December, this period encomp...
24/09/2026
Key Takeaways
Harmonic's cOS virtualized platform enables Lightcurve to de...
24/09/2026
COLD Launches Insider Club on Supercast, Giving True Crime Fans Ad-Free Listenin...
24/09/2026
The home interiors retailer signs up for third year as sponsor
RT Commercial today announced that Harry Corry has signed up to a third year as sponsor of Live...
24/09/2026
When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus' key proteins well eno...
24/09/2026
RT Radio 1 Folk Awards Tickets on Sale
Dolores Keane to be inducted into the...