Sony Pixel Power calrec Sony

NVIDIA Blackwell Raises Bar in New InferenceMAX Benchmarks, Delivering Unmatched Performance and Efficiency

09/10/2025

NVIDIA Blackwell swept the new SemiAnalysis InferenceMAX v1 benchmarks, delivering the highest performance and best overall efficiency.

InferenceMax v1 is the first independent benchmark to measure total cost of compute across diverse models and real-world scenarios.

Best return on investment: NVIDIA GB200 NVL72 delivers unmatched AI factory economics - a $5 million investment generates $75 million in DSR1 token revenue, a 15x return on investment.

Lowest total cost of ownership: NVIDIA B200 software optimizations achieve two cents per million tokens on gpt-oss, delivering 5x lower cost per token in just 2 months.

Best throughput and interactivity: NVIDIA B200 sets the pace with 60,000 tokens per second per GPU and 1,000 tokens per second per user on gpt-oss with the latest NVIDIA TensorRT-LLM stack.

As AI shifts from one-shot answers to complex reasoning, the demand for inference - and the economics behind it - is exploding.

The new independent InferenceMAX v1 benchmarks are the first to measure total cost of compute across real-world scenarios. The results? The NVIDIA Blackwell platform swept the field - delivering unmatched performance and best overall efficiency for AI factories.

A $5 million investment in an NVIDIA GB200 NVL72 system can generate $75 million in token revenue. That's a 15x return on investment (ROI) - the new economics of inference.

Inference is where AI delivers value every day, said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. These results show that NVIDIA's full-stack approach gives customers the performance and efficiency they need to deploy AI at scale.

Enter InferenceMAX v1 InferenceMAX v1, a new benchmark from SemiAnalysis released Monday, is the latest to highlight Blackwell's inference leadership. It runs popular models across leading platforms, measures performance for a wide range of use cases and publishes results anyone can verify.

Why do benchmarks like this matter?

Because modern AI isn't just about raw speed - it's about efficiency and economics at scale. As models shift from one-shot replies to multistep reasoning and tool use, they generate far more tokens per query, dramatically increasing compute demands.

NVIDIA's open-source collaborations with OpenAI (gpt-oss 120B), Meta (Llama 3 70B), and DeepSeek AI (DeepSeek R1) highlight how community-driven models are advancing state-of-the-art reasoning and efficiency.

Partnering with these leading model builders and the open-source community, NVIDIA ensures the latest models are optimized for the world's largest AI inference infrastructure. These efforts reflect a broader commitment to open ecosystems - where shared innovation accelerates progress for everyone.

Deep collaborations with the FlashInfer, SGLang and vLLM communities enable codeveloped kernel and runtime enhancements that power these models at scale.

Software Optimizations Deliver Continued Performance Gains NVIDIA continuously improves performance through hardware and software codesign optimizations. Initial gpt-oss-120b performance on an NVIDIA DGX Blackwell B200 system with the NVIDIA TensorRT LLM library was market-leading, but NVIDIA's teams and the community have significantly optimized TensorRT LLM for open-source large language models.

The TensorRT LLM v1.0 release is a major breakthrough in making large AI models faster and more responsive for everyone.

Through advanced parallelization techniques, it uses the B200 system and NVIDIA NVLink Switch's 1,800 GB/s bidirectional bandwidth to dramatically improve the performance of the gpt-oss-120b model.

The innovation doesn't stop there. The newly released gpt-oss-120b-Eagle3-v2 model introduces speculative decoding, a clever method that predicts multiple tokens at a time.

This reduces lag and delivers even quicker results, tripling throughput at 100 tokens per second per user (TPS/user) - boosting per-GPU speeds from 6,000 to 30,000 tokens.

For dense AI models like Llama 3.3 70B, which demand significant computational resources due to their large parameter count and the fact that all parameters are utilized simultaneously during inference, NVIDIA Blackwell B200 sets a new performance standard in InferenceMAX v1 benchmarks.

Blackwell delivers over 10,000 TPS per GPU at 50 TPS per user interactivity - 4x higher per-GPU throughput compared with the NVIDIA H200 GPU.

Performance Efficiency Drives Value Metrics like tokens per watt, cost per million tokens and TPS/user matter as much as throughput. In fact, for power-limited AI factories, Blackwell delivers 10x throughput per megawatt compared with the previous generation, which translates into higher token revenue.

The cost per token is crucial for evaluating AI model efficiency, directly impacting operational expenses. The NVIDIA Blackwell architecture lowered cost per million tokens by 15x versus the previous generation, leading to substantial savings and fostering wider AI deployment and innovation.

Multidimensional Performance InferenceMAX uses the Pareto frontier - a curve that shows the best trade-offs between different factors, such as data center throughput and responsiveness - to map performance.

But it's more than a chart. It reflects how NVIDIA Blackwell balances the full spectrum of production priorities: cost, energy efficiency, throughput and responsiveness. That balance enables the highest ROI across real-world workloads.

Systems that optimize for just one mode or scenario may show peak performance in isolation, but the economics of that doesn't scale. Blackwell's full-stack design delivers efficiency and value where it matters most: in production.

For a deeper look at how these curves are built - and why they matter for total cost of ownership and service-level agreement planning - check out this technical deep d
LINK: https://blogs.nvidia.com/blog/blackwell-inferencemax-benchmark-results...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

15/06/2026

Detach from Direct-Attached: How Remote Editing with EVO Keeps Creative Teams Moving

Detach from Direct-Attached: How Remote Editing with EVO Keeps Creative Teams Mo...

14/06/2026

Detroit Drums from Iconic Instruments

Library captures 1960s R&B/pop drum sound Following on from their recent wave of plug-in effects, Iconic Instruments have just launched an all-new virtual d...

14/06/2026

HBO Comedy Rooster Shot with URSA Cine 17K 65

HBO Comedy Rooster Shot with URSA Cine 17K 65 Brie Clayton June 14, 2026 0 Comments Large format brings viewers intimately close to characters. Black...

13/06/2026

Rhythmic Filters for Devious Machines' Infiltrator

Latest expansion pack includes 252 presets Devious Machines have recently introduced another expansion for their powerful multi-effects plug-in, Infiltrator...

13/06/2026

MetaGrid Pro gains AI Builder

Create custom DAW/plug-in controllers using prompts MetaGrid have recently introduced an all-new AI Builder function to their touchscreen-based control surf...

13/06/2026

Spectrum Reach Taps Anoki AI for Contextual Intelligence

Share Copy link Facebook X Linkedin Bluesky Email...

13/06/2026

Google TV Launches Soccer Hub, New Voice Command Features

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

YES Network and Gotham Sports App to Air Seven Athletes Unlimited Softball League Games

YES Network and The Gotham Sports App will air seven Athletes Unlimited Softball...

12/06/2026

UFL to Feature FAST Innovation Suite at 2026 United Bowl

The United Football League will host its FAST Innovation Suite at the 2026 United Bowl presented by Credit One Bank on Saturday, June 13 at 3:00 p.m. ET at Audi...

12/06/2026

InfoComm 2026: PTZOptics and LayerJot to Demo AI-Driven Camera Control

PTZOptics and LayerJot will present live demonstrations at InfoComm 2026 showing how natural-language AI prompting, robotic camera control, and on-device comput...

12/06/2026

InfoComm 2026: MultiDyne to Debut VF-9100 Fiber Transport Platform and Crescendo Audio Monitor

MultiDyne Video and Fiber Optic Systems will exhibit at InfoComm 2026, featuring...

12/06/2026

Eurovision Services Deploys Ateme Software-Based Frame-Rate Conversion

Ateme has announced that Eurovision Services is using Ateme's software-based frame-rate conversion technology for international live event workflows. The de...

12/06/2026

Bitmovin, Simplestream, and Xperi Partner to Support OTT Services on TiVo OS

Bitmovin and Simplestream have announced a partnership with Xperi to simplify the launch of OTT streaming services on TiVo OS smart TVs and devices. The collabo...

12/06/2026

Net Insight Deploys Nimbra 520 and Nimbra Edge for Multinational Corporate Live Production Workflow

Net Insight has announced that a multinational technology company is deploying a...

12/06/2026

MLB Players Inc., Athletes First Announce Content Partnership

MLB Players Inc., the business arm of the MLB Players Association, has announced a partnership with Athletes First to develop and sell brand partnerships across...

12/06/2026

G&D and VuWall Announce CommandKeyboard-Advanced for Network-Independent Control Room Operations

Guntermann and Drunck (G&D) and VuWall have announced the CommandKeyboard-Advanc...

12/06/2026

Philadelphia Union and Comcast Deploy Smart Technology at Subaru Park and WSFS Bank Sportsplex

Comcast Smart Solutions announces a new smart technology deployment with Major L...

12/06/2026

Elevation Worship Completes First Leg of 2026 Tour Using SSL Live Consoles and New UMD192 Interface

Elevation Worship completed the initial leg of its Elevation Nights 2026 tour ...

12/06/2026

AJA Announces KONA IP25 Integration with Colorfront Transkoder and On-Set Dailies

AJA Video Systems has announced KONA IP25 support for Colorfront Transkoder and ...

12/06/2026

InfoComm 2026: Audinate To Exhibit With New AVIO Install Adapters and Iris Camera Control Platform

Audinate Group Limited (ASX: AD8) will exhibit at InfoComm 2026 (Booth C7321, Ce...

12/06/2026

Pac-12 Appoints Scott Adametz as Chief Technology Officer

Pac-12 Commissioner Teresa Gould has announced the appointment of Scott Adametz as Chief Technology Officer. The Pac-12 describes the hire as the first CTO appo...

12/06/2026

InfoComm 2026: Grass Valley Introduces AMPP Edge Live for Enterprise Production

Grass Valley has announced AMPP Edge Live, a production system combining Grass Valley hardware, NVIDIA Blackwell GPU acceleration, and AMPP OS in a single platf...

12/06/2026

University of Texas's Brandon Rudy on a New Era of Live Sports Production in Austin

At one time a trailblazer with the launch of the Longhorn Network, the Universit...

12/06/2026

Ratings Roundup: NBA Finals Game 3 Hits 28-Year High; Stanley Cup Final Is Best Since 2015 Through Four Games

Ratings Roundup is a rundown of recent rating news and is derived from press rel...

12/06/2026

Chyron Releases PAINT 10.4 with Pro Football Data Integration and AI Player Cutout

Chyron has announced PAINT 10.4, an update to its illustrated replay and sports ...

12/06/2026

ESPN's MLB Productions Heat Up in June as Core Summer Schedule Gets Rolling

SVP, Production, Mark Gross: With the new schedule, with not having every Sunday night, it has given us an opportunity to take a step back and reimagine what o...

12/06/2026

Televisas IBC Team Delivers for Opening Mexico Match

For Televisa Technical Engineering Manager Roberto N nez Ibarra and the small team of 12 technicians and two production personnel at the IBC things are already ...

12/06/2026

GearExpo UK: Home Studio Acoustics Talk

Simple Steps to Better Acoustics - Taming The Small Room Most of us mix in spare rooms and small spaces, where the acoustics fight us at every turn. At Gear...

12/06/2026

Meris introduce the Ottobit X

Latest addition expands vintage-inspired effects palette Meris' Ottobit pedal range draws its inspiration from vintage gaming consoles, and the latest a...

12/06/2026

Sonora Cinematic release Movimento Strings Inflections

Soundbox-based chamber strings series expanded Sonora Cinematic have just announced the launch of the second instalment in their Soundbox-based chamber stri...

12/06/2026

Research: Mixed Picture for FIFA World Cup Broadcast Revenues

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

Viant Launches Enhanced Publisher Solutions for CTV, Programmatic

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

AJA Announces KONA IP25 Integration with Colorfront Software

AJA Announces KONA IP25 Integration with Colorfront Software Brie Clayton June 12, 2026 0 Comments Collaboration enables uncompressed SMPTE ST 2110 I/O ...

12/06/2026

URSA Cine 12K LF Used to Create Visuals for STUTS' K-Arena Concert

URSA Cine 12K LF Used to Create Visuals for STUTS' K-Arena Concert Brie Clayton June 12, 2026 0 Comments Organic visuals projected on a giant scre...

12/06/2026

MTI FILM Acquires Mango New Edit, Expanding its Global Post-Production Services From Set to Screen

MTI FILM Acquires Mango New Edit, Expanding its Global Post-Production Services ...

12/06/2026

AI Point Tracking Speeds Up Complex VFX Tracks in Mocha Pro

AI Point Tracking Speeds Up Complex VFX Tracks in Mocha Pro Jessie Electa Petrov June 12, 2026 0 Comments The 2026.5 release adds automatic point trac...

12/06/2026

Bitmovin Partners with Simplestream and Xperi to Support...

Bitmovin, a provider of video streaming solutions, has partnered with Simplestream, a provider of OTT and broadcast solutions, and technology provider Xperi, to...

12/06/2026

Jigsaw24 Signs Deal to Resell Leostream Remote Desktop Ac...

Leostream Corporation, creator of the world-leading Leostream Remote Desktop Access Platform, today announced Jigsaw24, a leading B2B IT solutions provider wit...

12/06/2026

Study: 2026 Election Cycle to Hit Record $11.6 Billion Ad Spend

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

NAB Elects Leadership at June Board of Directors Meeting

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

FCC Believes World Cup Communication Will Score Highly

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

Broadcasters Back NO FAKES Act

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

Scripps Unveils Coverage Plans For America's 250th Anniversary

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

How Aussie indie games and screen are levelling up with IP

How Aussie indie games and screen are levelling up with IP 11 June 2026 Ari Harrison, Pro Jank Footy Head of Games Joey Egger and Ari Harrison of Umbrella sha...

12/06/2026

NVIDIA Blackwell Leads on First Agentic AI Infrastructure Benchmark

AgentPerf from Artificial Analysis, the industry's first agentic AI benchmark, gives developers, enterprises and infrastructure providers a clear way to com...