
NVIDIA Blackwell swept the new SemiAnalysis InferenceMAX v1 benchmarks, delivering the highest performance and best overall efficiency.
InferenceMax v1 is the first independent benchmark to measure total cost of compute across diverse models and real-world scenarios.
Best return on investment: NVIDIA GB200 NVL72 delivers unmatched AI factory economics - a $5 million investment generates $75 million in DSR1 token revenue, a 15x return on investment.
Lowest total cost of ownership: NVIDIA B200 software optimizations achieve two cents per million tokens on gpt-oss, delivering 5x lower cost per token in just 2 months.
Best throughput and interactivity: NVIDIA B200 sets the pace with 60,000 tokens per second per GPU and 1,000 tokens per second per user on gpt-oss with the latest NVIDIA TensorRT-LLM stack.
As AI shifts from one-shot answers to complex reasoning, the demand for inference - and the economics behind it - is exploding.
The new independent InferenceMAX v1 benchmarks are the first to measure total cost of compute across real-world scenarios. The results? The NVIDIA Blackwell platform swept the field - delivering unmatched performance and best overall efficiency for AI factories.
A $5 million investment in an NVIDIA GB200 NVL72 system can generate $75 million in token revenue. That's a 15x return on investment (ROI) - the new economics of inference.
Inference is where AI delivers value every day, said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. These results show that NVIDIA's full-stack approach gives customers the performance and efficiency they need to deploy AI at scale.
Enter InferenceMAX v1 InferenceMAX v1, a new benchmark from SemiAnalysis released Monday, is the latest to highlight Blackwell's inference leadership. It runs popular models across leading platforms, measures performance for a wide range of use cases and publishes results anyone can verify.
Why do benchmarks like this matter?
Because modern AI isn't just about raw speed - it's about efficiency and economics at scale. As models shift from one-shot replies to multistep reasoning and tool use, they generate far more tokens per query, dramatically increasing compute demands.
NVIDIA's open-source collaborations with OpenAI (gpt-oss 120B), Meta (Llama 3 70B), and DeepSeek AI (DeepSeek R1) highlight how community-driven models are advancing state-of-the-art reasoning and efficiency.
Partnering with these leading model builders and the open-source community, NVIDIA ensures the latest models are optimized for the world's largest AI inference infrastructure. These efforts reflect a broader commitment to open ecosystems - where shared innovation accelerates progress for everyone.
Deep collaborations with the FlashInfer, SGLang and vLLM communities enable codeveloped kernel and runtime enhancements that power these models at scale.
Software Optimizations Deliver Continued Performance Gains NVIDIA continuously improves performance through hardware and software codesign optimizations. Initial gpt-oss-120b performance on an NVIDIA DGX Blackwell B200 system with the NVIDIA TensorRT LLM library was market-leading, but NVIDIA's teams and the community have significantly optimized TensorRT LLM for open-source large language models.
The TensorRT LLM v1.0 release is a major breakthrough in making large AI models faster and more responsive for everyone.
Through advanced parallelization techniques, it uses the B200 system and NVIDIA NVLink Switch's 1,800 GB/s bidirectional bandwidth to dramatically improve the performance of the gpt-oss-120b model.
The innovation doesn't stop there. The newly released gpt-oss-120b-Eagle3-v2 model introduces speculative decoding, a clever method that predicts multiple tokens at a time.
This reduces lag and delivers even quicker results, tripling throughput at 100 tokens per second per user (TPS/user) - boosting per-GPU speeds from 6,000 to 30,000 tokens.
For dense AI models like Llama 3.3 70B, which demand significant computational resources due to their large parameter count and the fact that all parameters are utilized simultaneously during inference, NVIDIA Blackwell B200 sets a new performance standard in InferenceMAX v1 benchmarks.
Blackwell delivers over 10,000 TPS per GPU at 50 TPS per user interactivity - 4x higher per-GPU throughput compared with the NVIDIA H200 GPU.
Performance Efficiency Drives Value Metrics like tokens per watt, cost per million tokens and TPS/user matter as much as throughput. In fact, for power-limited AI factories, Blackwell delivers 10x throughput per megawatt compared with the previous generation, which translates into higher token revenue.
The cost per token is crucial for evaluating AI model efficiency, directly impacting operational expenses. The NVIDIA Blackwell architecture lowered cost per million tokens by 15x versus the previous generation, leading to substantial savings and fostering wider AI deployment and innovation.
Multidimensional Performance InferenceMAX uses the Pareto frontier - a curve that shows the best trade-offs between different factors, such as data center throughput and responsiveness - to map performance.
But it's more than a chart. It reflects how NVIDIA Blackwell balances the full spectrum of production priorities: cost, energy efficiency, throughput and responsiveness. That balance enables the highest ROI across real-world workloads.
Systems that optimize for just one mode or scenario may show peak performance in isolation, but the economics of that doesn't scale. Blackwell's full-stack design delivers efficiency and value where it matters most: in production.
For a deeper look at how these curves are built - and why they matter for total cost of ownership and service-level agreement planning - check out this technical deep d
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
29/04/2026
It was a delicate job in a 150-year-old venue laden with traditions. Begun at th...
29/04/2026
In annual event, the league gives startup companies the opportunity to prove the...
29/04/2026
Panel discussions, networking, and a facility tour will take place in the renova...
29/04/2026
(L-R) Derek Drescher, Coss Marte, and Syretta Wright have each other's backs. (Micheal Hurcomb/Shutterstock for Sundance Film Festival)
By Veronika Lee Cla...
29/04/2026
Combines EQ and harmonic distortion
Techivation's latest release is a simple EQ designed to offer quick control over a source's overall tonal balanc...
29/04/2026
Two new MPE controllers announced
Expressive E caused quite a stir when they released the Osmose, making the sort of expression that was once reserved for p...
29/04/2026
New modules & enhanced machine-learning
The latest version of iZotope's flagship restoration suite is now available, and now offers over 50 tools design...
29/04/2026
Surgeon Dr Jasmina Kevric wins 2026 Les Murray Award
29 April, 2026
Media releases
Australia for UNHCR and SBS are proud to announce that Dr Jasmina Kevric...
29/04/2026
Some people stumble into their passion. Julissa Padilla walked straight into a film vault. For her, entertainment was never just about the movies themselves. It...
29/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
29/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
29/04/2026
Clear-Com has appointed Brian Grahn as Market Outreach Manager of the Americas and Ben Turnwell as Business Development Manager for EMEA live, expanding their ...
29/04/2026
nxtedition is bringing its range of consolidated production tools to MPTS 2026, with new developments spanning transcription, editing, graphics and AI-assisted ...
29/04/2026
Quortex Switch to boost the streaming experience for Telxius customers, reaching millions of viewers worldwide
Synamedia and Telxius, the leading global connec...
29/04/2026
freispace, the leading ERP-as-a-Service platform for media and entertainment production, and Projective, a leading provider of post-production collaboration tec...
29/04/2026
DHD reports strong interest in its broadcast audio product range, exhibited at the April 19th-22nd NAB Show in Las Vegas. The event attracted a claimed 58,000 a...
29/04/2026
Student Spotlight: Alan Catz The Argentine film and game composer talks about working on League of Legends, receiving Berklee's BMI Award, and the lifelon...
29/04/2026
Jay Jennings Builds the Worlds You Hear on Screen The supervising sound designer behind A Minecraft Movie, The Meg, Letters from Iwo Jima, and dozens of other...
29/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
29/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
29/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
29/04/2026
29 Apr 2026
VEON and Kyivstar Fulfill Commitment to Invest USD 1 Billion in Ukr...
29/04/2026
Rhod Gilbert, Harriet Kemsley, Kae Kurd, Sara Pascoe and Vicki Pattison to take part in brand new series on free streaming service U
London, 29th April 2026: F...
29/04/2026
Wednesday 29 April 2026
Katie Price: Nothing to Hide, a Sky Original documentar...
29/04/2026
Re-examining the case of Ellie Williams and the wider story of grooming in the town of BarrowWednesday 29 April 2026
Sky announces upcoming documentary series ...
29/04/2026
Wednesday 29 April 2026
Jennifer Garner to lead an all-star cast in new Sky Exc...
29/04/2026
Back to All News
SUPERNOVA: GENESIS Reached a Peak Audience of More Than 6.5 Mi...
29/04/2026
Students and staff from Hills Road Sixth Form College in Cambridge ran a 4.5km course around the roads of Cambridge as part of their annual programme of sustain...
29/04/2026
The Dawn Chorus airs Sunday 3 May from midnight to 7am on RT Radio 1 and RT ly...
29/04/2026
Jin-Quan Yu elected to the National Academy of Sciences Yu is recognized for his pioneering work in synthetic organic chemistry.
April 28, 2026
LA JOLLA, CA S...
28/04/2026
The audio team for the entertainment event must blend speech intelligibility with full-range music reproduction while considering the broadcast
Last week's...
28/04/2026
The Pac-12 Conference has released an updated primary mark and logo as the starting point of the new league's brand identity. The mark was soft-launched acr...
28/04/2026
The DP World Tour and Amazon Leo have signed an agreement making Amazon's lo...
28/04/2026
Pixellot and HELIOS have announced an integration that automatically converts full-game hockey video into individualized shift videos for each athlete, without ...
28/04/2026
Daktronics has partnered with the Asheville Tourists to manufacture and install a new LED video display. The installation was completed in late 2025 and is now ...
28/04/2026
Eutelsat has announced the renewal of its partnership with PCTV, a content aggregation and distribution company in Mexico and part of Megacable Holdings, for co...
28/04/2026
Daktronics has partnered with the Gary SouthShore RailCats to install a new LED video display at U.S. Steel Yard, replacing the previous Daktronics display inst...
28/04/2026
Telos Alliance and the College Radio Foundation have announced that WWSU-FM of W...
28/04/2026
Golf viewership is growing. The 2025 Ryder Cup drew five million viewers in the UK, a 45% increase over the 2023 event. The US Open was the most streamed golf e...
28/04/2026
The CW Network and WWE, part of TKO Group Holdings (NYSE: TKO), have announced t...
28/04/2026
The Alliance for IP Media Solutions (AIMS) has announced that the Internet Protocol Media Experience (IPMX) suite of standards and specifications has been named...
28/04/2026
The 2026 NAB Show is in the books and the show once again served up a cavalcade ...
28/04/2026
Gray Media and RAJ Sports have announced Rose City SportsNet (RCSN), a new netwo...
28/04/2026
Today, we announced our First Quarter 2026 earnings, starting the Year of Raising Ambition with strong momentum across the business and continued innovation acr...