Sony Pixel Power calrec Sony

Leading Inference Providers Cut AI Costs by up to 10x With Open Source Models on NVIDIA Blackwell

12/02/2026

A diagnostic insight in healthcare. A character's dialogue in an interactive game. An autonomous resolution from a customer service agent. Each of these AI-powered interactions is built on the same unit of intelligence: a token.

Scaling these AI interactions requires businesses to consider whether they can afford more tokens. The answer lies in better tokenomics - which at its core is about driving down the cost of each token. This downward trend is unfolding across industries. Recent MIT research found that infrastructure and algorithmic efficiencies are reducing inference costs for frontier-level performance by up to 10x annually.

To understand how infrastructure efficiency improves tokenomics, consider the analogy of a high-speed printing press. If the press produces 10x output with incremental investment in ink, energy and the machine itself, the cost to print each individual page drops. In the same way, investments in AI infrastructure can lead to far greater token output compared with the increase in cost - causing a meaningful reduction in the cost per token.

When token output outpaces infrastructure cost, the cost of each token drops. That's why leading inference providers including Baseten, DeepInfra, Fireworks AI and Together AI are using the NVIDIA Blackwell platform, which helps them reduce cost per token by up to 10x compared with the NVIDIA Hopper platform.

These providers host advanced open source models, which have now reached frontier-level intelligence. By combining open source frontier intelligence, the extreme hardware-software codesign of NVIDIA Blackwell and their own optimized inference stacks, these providers are enabling dramatic token cost reductions for businesses across every industry.

Healthcare - Baseten and Sully.ai Cut AI Inference Costs by 10x In healthcare, tedious, time-consuming tasks like medical coding, documentation and managing insurance forms cut into the time doctors can spend with patients.

Sully.ai helps solve this problem by developing AI employees that can handle routine tasks like medical coding and note-taking. As the company's platform scaled, its proprietary, closed source models created three bottlenecks: unpredictable latency in real-time clinical workflows, inference costs that scaled faster than revenue and insufficient control over model quality and updates.

Sully.ai builds AI employees that handle routine tasks for physicians. To overcome these bottlenecks, Sully.ai uses Baseten's Model API, which deploys open source models such as gpt-oss-120b on NVIDIA Blackwell GPUs. Baseten used the low-precision NVFP4 data format, the NVIDIA TensorRT-LLM library and the NVIDIA Dynamo inference framework to deliver optimized inference. The company chose NVIDIA Blackwell to run its Model API after seeing up to 2.5x better throughput per dollar compared with the NVIDIA Hopper platform.

As a result, Sully.ai's inference costs dropped by 90%, representing a 10x reduction compared with the prior closed source implementation, while response times improved by 65% for critical workflows like generating medical notes. The company has now returned over 30 million minutes to physicians, time previously lost to data entry and other manual tasks.

Gaming - DeepInfra and Latitude Reduce Cost per Token by 4x Latitude is building the future of AI-native gaming with its AI Dungeon adventure-story game and upcoming AI-powered role-playing gaming platform, Voyage, where players can create or play worlds with the freedom to choose any action and make their own story.

The company's platform uses large language models to respond to players' actions - but this comes with scaling challenges, as every player action triggers an inference request. Costs scale with engagement, and response times must stay fast enough to keep the experience seamless.

Latitude has built a text-based adventure-story game called AI Dungeon, which generates both narrative text and imagery in real time as players explore dynamic stories. Latitude runs large open source models on DeepInfra's inference platform, powered by NVIDIA Blackwell GPUs and TensorRT-LLM. For a large-scale mixture-of-experts (MoE) model, DeepInfra reduced the cost per million tokens from 20 cents on the NVIDIA Hopper platform to 10 cents on Blackwell. Moving to Blackwell's native low-precision NVFP4 format further cut that cost to just 5 cents - for a total 4x improvement in cost per token - while maintaining the accuracy that customers expect.

Running these large-scale MoE models on DeepInfra's Blackwell-powered platform allows Latitude to deliver fast, reliable responses cost effectively. DeepInfra inference platform delivers this performance while reliably handling traffic spikes, letting Latitude deploy more capable models without compromising player experience.

Agentic Chat - Fireworks AI and Sentient Foundation Lower AI Costs by up to 50% Sentient Labs is focused on bringing AI developers together to build powerful reasoning AI systems that are all open source. The goal is to accelerate AI toward solving harder reasoning problems through research in secure autonomy, agentic architecture and continual learning.

Its first app, Sentient Chat, orchestrates complex multi-agent workflows and integrates more than a dozen specialized AI agents from the community. Due to this, Sentient Chat has massive compute demands because a single user query could trigger a cascade of autonomous interactions that typically lead to costly infrastructure overhead.

To manage this scale and complexity, Sentient uses Fireworks AI's inference platform running on NVIDIA Blackwell. With Fireworks' Blackwell-optimized inference stack, Sentient achieved 25-50% better cost efficiency compared with its previous Hopper-based deployment.

Sentient Chat orchestrates complex multi-agent work
LINK: https://blogs.nvidia.com/blog/inference-open-source-models-blackwell-r...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

01/04/2026

DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION

January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION Douyin Users Can Now Create And Share Videos With Stun...

12/02/2026

LTN Names 3 Executives to Lead Technology Group

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

On the Ice, There's a Third Team at Work

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Mo Rocca to Receive the 2026 LABF Insight Award at NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Marshall Electronics POVs Power Hidden Camera Investigat...

The production team of the long-running German investigative series Achtung Abzocke recently upgraded its cameras for the show's 12th season. The objectiv...

12/02/2026

Bitmovin Appoints Ian Baglow as Co-Chief Executive Office...

Leading provider of video streaming solutions, Bitmovin, has appointed Ian Baglow as Co-CEO alongside existing CEO and Co-Founder Stefan Lederer. Under this str...

12/02/2026

Vizrt Launches Sports Production Bundles to Empower US St...

Vizrt, a leading viewer engagement platform and a trusted expert in live production technologies, today announces the launch of four Campus Stadium Production B...

12/02/2026

Ailanto and Cubbit launch sovereign cloud storage for Swi...

Strategic agreement to deliver S3 cloud storage in Switzerland with full data sovereignty and local control including at the level of individual cantons plu...

12/02/2026

Mad About Video counts on Lightware MX2 matrix switcher a...

Mad About Video is a leading specialist in video for live events and installations throughout Malta. In operation since 2011, it has evolved from a company focu...

12/02/2026

JAGGAER supports Betsson Group in further strengthening i...

JAGGAER, a global leader in digital procurement and supplier collaboration solutions, today announced the successful delivery of a procurement digitalization pr...

12/02/2026

LiveU Spotlights Three Broadcast Priorities at NAB Show 2...

At NAB Show, LiveU will showcase its broadest IP-video EcoSystem to date, designed to help broadcasters and content creators embrace digital first operations, d...

12/02/2026

Spectrum News Acquires New England Cable News

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Vizrt Unveils Campus Stadium Production Bundles

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Hulu + Live TV Adds Fubo Sports Network to Channel Line-up

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

FCC To Hold Open Commission Meeting on Feb. 18

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Ralph M. Oakley to Receive NAB's Chuck Sherman TV Leadership Award

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

NVIDIA DGX Spark Powers Big Projects in Higher Education

At leading institutions across the globe, the NVIDIA DGX Spark desktop supercomputer is bringing data center class AI to lab benches, faculty offices and studen...

12/02/2026

Leading Inference Providers Cut AI Costs by up to 10x With Open Source Models on NVIDIA Blackwell

A diagnostic insight in healthcare. A character's dialogue in an interactive...

12/02/2026

GeForce NOW Turns Screens Into a Gaming Machine

The GeForce NOW sixth-anniversary festivities roll on this February, continuing a monthlong celebration of NVIDIA's cloud gaming service. This week brings ...

12/02/2026

February 11, 2026

TIME100 Health list features Scripps Research Professor Darrell Irvine Irvine is recognized for his work in empowering the immune system to fight disease, which...

11/02/2026

FYI: Phone Support Maintenance

FYI: Phone Support Maintenance One thing we pride ourselves on here at Utah Scientific is our 24-hour support included with our signature 10-year hardware warra...

11/02/2026

Bitmovin Appoints Ian Baglow as Co-Chief Executive Officer

Leading provider of video streaming solutions, Bitmovin, has appointed Ian Baglow as Co-CEO alongside existing CEO and Co-Founder Stefan Lederer. Under this str...

11/02/2026

Paramount and CBS Partner to Air UFC 326

Paramount and the CBS Television Network will partner to air UFC 326: HOLLOWAY vs. OLIVEIRA 2 live on Saturday, March 7, from T-Mobile Arena in Las Vegas, mar...

11/02/2026

MLB.TV Launches on ESPN Beginning February 10

Beginning February 10, fans can buy MLB.TV on ESPN, a new milestone in one of sports media's longest-standing partnerships. ESPN becomes the new streaming h...

11/02/2026

Fubo Sports Network Launches on Hulu + Live TV

Fubo Sports Network is available to Hulu's Live TV subscribers in the core $89.99 a month subscription plan, which also includes full access to the entire H...

11/02/2026

Rai Selects Imagine Selenio Network Processor for IP Migration

Following a competitive public tender process, Rai (Radiotelevisione Italiana), the national public broadcasting company of Italy, has awarded Imagine Communica...

11/02/2026

MLB Makes In-Market Streaming Subscriptions for 20 Clubs Available to Fans

Major League Baseball is making in-market streaming subscriptions for 20 Clubs available today for fans. Subscriptions for the following Clubs are available vi...

11/02/2026

5G Broadcast Trials Return to the Olympic Stage at Milano Cortina 2026

Building on successful demonstrations during the Paris Olympics 2024, Italian public service broadcaster Rai and the European Broadcasting Union (EBU) are condu...

11/02/2026

ESPN and Disney Launch We're Going, the First Marketing Campaign for ESPN's Inaugural Super Bowl

Following Sunday's Super Bowl LX, ESPN and Disney unveiled We're Going,...

11/02/2026

Stats Perform: 2026 Super Bowl Latency Report

Delayed streams are a growing source of frustration for sports fans. During the 2026 Super Bowl, some streams lagged up to 62 seconds behind the action on the f...

11/02/2026

NASCAR Channel and FloSports to Simulcast 16 Races Live

NASCAR and FloSports announces an expanded slate of racing events that will bring FloRacing coverage live throughout the 2026 season to the NASCAR Channel, furt...

11/02/2026

Manifold Expands Sales Presence in Europe

Manifold technologies GmbH announces the appointment of Nick Tucker as Sales Manager for Europe, reinforcing the company's continued growth across broadcast...

11/02/2026

Genies and MLB Players Inc. Team Up to Create AI Characters of MLB Players

Genies, the AI avatar technology company powering the next era of interactive digital identity, entered into a landmark collaboration with MLB Players, Inc., th...

11/02/2026

ICC, Google Partner for the First-Ever AI-Powered ICC Men's T20 World Cup fuelled by Gemini & Pixel

The International Cricket Council (ICC) and Google have joined forces for an AI-...

11/02/2026

Dolby Highlights From First Ever Super Bowl LX Innovation Summit

Dolby's CEO Kevin Yeaman and Giles Baker, SVP of Dolby Cloud Solutions, shared how the brand's latest innovations - Dolby Vision, Dolby Atmos, and Dolby...

11/02/2026

Detroit Tigers and Detroit Red Wings Enter Broadcast Partnership with MLB

Ilitch Sports + Entertainment has entered a first of its kind partnership with Major League Baseball, which will provide broadcast support to both the Detroit T...

11/02/2026

A Changing Landscape: MLB Local Media Brings Detroit Tigers, Los Angeles Angels Into the Fold; ESPN Begins Distribution of MLB.TV

Broadcasts of the NHL's Detroit Red Wings will also be produced by the leagu...

11/02/2026

DAM Los Angeles

Video moves fast can your DAM keep up? Join Blue Lucy in LA for the West Coast's leading Digital Asset Management event as we explore, celebrate, and acc...

11/02/2026

Super Bowl LX Delivers 124.9 Million Viewers

NEW YORK - February 10, 2026 - An estimated 124.9 million viewers watched Super Bowl LX on Sunday, February 8, according to Nielsen's Big Data Panel measu...

11/02/2026

Scripps Selling Court TV to Jellysmack

Share Copy link Facebook X Linkedin Bluesky Email...

11/02/2026

Schulze-Brakel Introduces Wireless Clip-On Branding For Mics

Share Copy link Facebook X Linkedin Bluesky Email...

11/02/2026

LiveU To Showcase Expanded IP-Video EcoSystem at 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

11/02/2026

Clear-Com Powers TEDNext 2025 with Gen-IC Virtual Interco...

Clear-Com provided an advanced, IP-based communications infrastructure for TEDNext 2025, supporting production, media, and editorial teams with a highly flexib...

11/02/2026

Astera Releases QuikBeam - Versatile 200W Equivalent LED...

Astera introduces QuikBeam, the newest addition to its acclaimed Quik family of focusing LED Fresnels. This ultra-compact spotlight combines the equivalent powe...

11/02/2026

Rai Selects Imagine Selenio Network Processor for IP Migr...

Following a competitive public tender process, Rai (Radiotelevisione Italiana), the national public broadcasting company of Italy, has awarded Imagine Communica...