
Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and to build entirely new classes of applications.
But with opportunities often come challenges.
Both on premises and in the cloud, applications that are expected to run in real time place significant demands on data center infrastructure to simultaneously deliver high throughput and low latency with one platform investment.
To drive continuous performance improvements and improve the return on infrastructure investments, NVIDIA regularly optimizes the state-of-the-art community models, including Meta's Llama, Google's Gemma, Microsoft's Phi and our own NVLM-D-72B, released just a few weeks ago.
Relentless Improvements Performance improvements let our customers and partners serve more complex models and reduce the needed infrastructure to host them. NVIDIA optimizes performance at every layer of the technology stack, including TensorRT-LLM, a purpose-built library to deliver state-of-the-art performance on the latest LLMs. With improvements to the open-source Llama 70B model, which delivers very high accuracy, we've already improved minimum latency performance by 3.5x in less than a year.
We're constantly improving our platform performance and regularly publish performance updates. Each week, improvements to NVIDIA software libraries are published, allowing customers to get more from the very same GPUs. For example, in just a few months' time, we've improved our low-latency Llama 70B performance by 3.5x.
NVIDIA has increased performance on the Llama 70B model by 3.5x. In the most recent round of MLPerf Inference 4.1, we made our first-ever submission with the Blackwell platform. It delivered 4x more performance than the previous generation.
This submission was also the first-ever MLPerf submission to use FP4 precision. Narrower precision formats, like FP4, reduces memory footprint and memory traffic, and also boost computational throughput. The process takes advantage of Blackwell's second-generation Transformer Engine, and with advanced quantization techniques that are part of TensorRT Model Optimizer, the Blackwell submission met the strict accuracy targets of the MLPerf benchmark.
Blackwell B200 delivers up to 4x more performance versus previous generation on MLPerf Inference v4.1's Llama 2 70B workload. Improvements in Blackwell haven't stopped the continued acceleration of Hopper. In the last year, Hopper performance has increased 3.4x in MLPerf on H100 thanks to regular software advancements. This means that NVIDIA's peak performance today, on Blackwell, is 10x faster than it was just one year ago on Hopper.
These results track progress on the MLPerf Inference Llama 2 70B Offline scenario over the past year. Our ongoing work is incorporated into TensorRT-LLM, a purpose-built library to accelerate LLMs that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM is built on top of the TensorRT Deep Learning Inference library and leverages much of TensorRT's deep learning optimizations with additional LLM-specific improvements.
Improving Llama in Leaps and Bounds More recently, we've continued optimizing variants of Meta's Llama models, including versions 3.1 and 3.2 as well as model sizes 70B and the biggest model, 405B. These optimizations include custom quantization recipes, as well as efficient use of parallelization techniques to more efficiently split the model across multiple GPUs, leveraging NVIDIA NVLink and NVSwitch interconnect technologies. Cutting-edge LLMs like Llama 3.1 405B are very demanding and require the combined performance of multiple state-of-the-art GPUs for fast responses.
Parallelism techniques require a hardware platform with a robust GPU-to-GPU interconnect fabric to get maximum performance and avoid communication bottlenecks. Each NVIDIA H200 Tensor Core GPU features fourth-generation NVLink, which provides a whopping 900GB/s of GPU-to-GPU bandwidth. Every eight-GPU HGX H200 platform also ships with four NVLink Switches, enabling every H200 GPU to communicate with any other H200 GPU at 900GB/s, simultaneously.
Many LLM deployments use parallelism over choosing to keep the workload on a single GPU, which can have compute bottlenecks. LLMs seek to balance low latency and high throughput, with the optimal parallelization technique depending on application requirements.
For instance, if lowest latency is the priority, tensor parallelism is critical, as the combined compute performance of multiple GPUs can be used to serve tokens to users more quickly. However, for use cases where peak throughput across all users is prioritized, pipeline parallelism can efficiently boost overall server throughput.
The table below shows that tensor parallelism can deliver over 5x more throughput in minimum latency scenarios, whereas pipeline parallelism brings 50% more performance for maximum throughput use cases.
For production deployments that seek to maximize throughput within a given latency budget, a platform needs to provide the ability to effectively combine both techniques like in TensorRT-LLM.
Read the technical blog on boosting Llama 3.1 405B throughput to learn more about these techniques.
Different scenarios have different requirements, and parallelism techniques bring optimal performance for each of these scenarios. The Virtuous Cycle Over the lifecycle of our architectures, we deliver significant performance gains from ongoing software tuning and optimization. These improvements translate into additional value for customers who train and deploy on our platforms. They're able to create more capable models and applications and deploy their existing models using less infrastructure, enhancing th
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
04/09/2026
Eluvio will demonstrate updates to its Content Fabric platform at IBC2026 (Hall 8, Stand 8.MS5, and Future Tech Zone, Hall 14, Stand 14.A58, September 11-14). C...
04/09/2026
LiveU has joined the IBC Accelerator's Smart Stories consortium, a multi-vendor initiative developing an open standard for Story Context interoperability in...
04/09/2026
Telos Alliance will demonstrate new audio observability capabilities at IBC2026 (Stand 8.D37, RAI Amsterdam, September 11-14), covering monitoring of audio heal...
04/09/2026
World Team Tennis (WTT) has announced a broadcast partnership with USA Sports, the sports division of Versant. USA Sports will air all six of WTT's live mat...
04/09/2026
Gravity Media has appointed Alex Pannell as Chief Business Officer. Pannell joins from Appear, where he served as Chief Commercial Officer. He will be based in ...
04/09/2026
Dalet has announced new agentic workflow capabilities for Dalia, its AI platform for media organizations, ahead of IBC2026. The updates extend Dalia from an AI ...
04/09/2026
L-Acoustics has achieved EN 54 certification across a portfolio of loudspeakers,...
04/09/2026
EVS and Matrox Video have announced a technology partnership integrating Matrox ...
04/09/2026
NEP Group will exhibit at IBC2026 (Hall 7, Stand 7.B55, RAI Amsterdam, September 11-14), demonstrating software-defined production capabilities across NEP Platf...
04/09/2026
The Television Academy has awarded an Engineering Emmy to Dr. Roman Foltyn, Chief Engineer of FoMaSystems, and Curt O. Schaller for the development of the ARRI ...
04/09/2026
InSync Technology will exhibit at IBC2026 (Stand 1.A55, RAI Amsterdam, September 11-14), demonstrating its FrameFormer and NOVA conversion platforms for frame r...
04/09/2026
The Professional Triathletes Organisation (PTO) has signed a multi-year agreemen...
04/09/2026
Accedo and MediaKind have announced a white-label, end-to-end streaming solution for operators, combining Accedo's native application framework and cloud-ba...
04/09/2026
TAG Video Systems has announced TAG Blue, a cloud platform for shared monitoring visibility across video delivery chains, with AI-powered root-cause analysis. T...
04/09/2026
TwelveLabs has launched a video compliance platform for media organizations, inc...
04/09/2026
The Television Academy and the National Academy of Television Arts & Sciences ha...
04/09/2026
Powerful new soft synth announced
Traveler is the first software instrument to be released by KeySolutions Sounds, and it's an ambitious one! Said to co...
04/09/2026
Free virtual instrument range expanded
The latest expansion of e-instruments' free Fragments series introduces a packed-down version of the company'...
04/09/2026
Company recognised for second time
Two years after winning an Engineering, Science & Technology Emmy Award for dxRevive Pro, Accentize have been recognised ...
04/09/2026
Bilbao - 3rd September 2026 -AgileTV, a leading provider of end-to-end TV and vi...
04/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/09/2026
ACEMAGIC First to Introduce Ryzen AI Max PRO 495 AI Mini Workstation at IFA 202...
04/09/2026
Krotos Brings Video to Sound Footsteps Directly into Premiere Pro and DaVinci Re...
04/09/2026
RED DIGITAL CINEMA Comes Out Swinging at IBC 2026 with Next-Generation Live Broa...
04/09/2026
Orty v2.0-web - the cold-outreach assistant moves to the web
Brie Clayton September 4, 2026
0 Comments
Orty today announced v2.0-web of its AI cold-ou...
04/09/2026
Lightware is expanding its LARA (Lightware Advanced Room Automation) platform with the GUDE Power Distribution Units (PDU) driver module. GUDE devices can now b...
04/09/2026
Lightware shares the second release in its ESG communications series. Following the first piece, which focused on the companys environmental performance, this r...
04/09/2026
Big Blue Marble supported a leading European public service broadcaster in delivering the Soccer World Cup 2026 streams, supplying more than 10 Tbps of bandwidt...
04/09/2026
Outstanding achievement in engineering, science and technology
Hitomi Broadcast has been recognised with an Engineering, Science & Technology Emmy from the Ac...
04/09/2026
California, Hollywood at a Crossroads
Andy Marken September 3, 2026
0 Comments
I'm doing everything alone. Crossing roads alone. Seeing the Eiffe...
04/09/2026
Berklee Announces Winners of 2026 SongwritersdB Songwriting Contest Students Luca Alexandru, Annika Wild, and Johanna James received the top prizes for origin...
04/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/09/2026
RT ANNOUNCE NEW DRAMA THE MURDER OF MR. MOONLIGHT
BASED ON THE DISAPPEARANCE...
03/09/2026
Yospace will be at IBC from 11th - 14th September (booth 5.C78). For more inform...
03/09/2026
NEP Group has announced leadership changes at Creative Technology Group. Founder and CEO Graham Andrews will step away from day-to-day leadership at the end of ...
03/09/2026
The NFL and TikTok have announced a multi-year partnership renewal expanding the league's presence on the platform. NFL posts on TikTok increased more than ...
03/09/2026
Grass Valley has appointed Ronny Van Geel as Vice President, Cameras. Van Geel began his career in 1999 at Philips DVS and moved into product management in 2005...
03/09/2026
Telestream has announced Vantage Super Resolution, an AI-powered upscaling capability for the Vantage media workflow platform, developed in collaboration with N...
03/09/2026
Solid State Logic (SSL) will debut the TCA Tour at IBC2026 (Hall 8, Stand 8.B45), a portable production system built from System T components in a desktop frame...
03/09/2026
DPA Microphones has announced the 6388 CORE Headset Microphone, making its debut at IBC2026 (Hall 8, Stand D60). The supercardioid headset features a physicall...
03/09/2026
Qvest will present a multi-vendor Dynamic Media Facility (DMF) showcase at IBC2026 (Booth 10.C24), demonstrating how the EBU's DMF architecture can be imple...
03/09/2026
Imagine Communications will showcase the Magellan PanelFlex control surface and other control portfolio updates at IBC2026 (Stand 1.B73, RAI Amsterdam, Septembe...
03/09/2026
Bridge Technologies will announce at IBC2026 (Booth 1.A71) that its VB440 production probe is now available as a software-based solution, deployable on commerci...
03/09/2026
Audio-Technica has unveiled the BP350ST MS (mid-side) microphone in two versions: the camera-mount BP350ST-UL and the boundary or gooseneck-mountable BP350ST-UB...