
A diagnostic insight in healthcare. A character's dialogue in an interactive game. An autonomous resolution from a customer service agent. Each of these AI-powered interactions is built on the same unit of intelligence: a token.
Scaling these AI interactions requires businesses to consider whether they can afford more tokens. The answer lies in better tokenomics - which at its core is about driving down the cost of each token. This downward trend is unfolding across industries. Recent MIT research found that infrastructure and algorithmic efficiencies are reducing inference costs for frontier-level performance by up to 10x annually.
To understand how infrastructure efficiency improves tokenomics, consider the analogy of a high-speed printing press. If the press produces 10x output with incremental investment in ink, energy and the machine itself, the cost to print each individual page drops. In the same way, investments in AI infrastructure can lead to far greater token output compared with the increase in cost - causing a meaningful reduction in the cost per token.
When token output outpaces infrastructure cost, the cost of each token drops. That's why leading inference providers including Baseten, DeepInfra, Fireworks AI and Together AI are using the NVIDIA Blackwell platform, which helps them reduce cost per token by up to 10x compared with the NVIDIA Hopper platform.
These providers host advanced open source models, which have now reached frontier-level intelligence. By combining open source frontier intelligence, the extreme hardware-software codesign of NVIDIA Blackwell and their own optimized inference stacks, these providers are enabling dramatic token cost reductions for businesses across every industry.
Healthcare - Baseten and Sully.ai Cut AI Inference Costs by 10x In healthcare, tedious, time-consuming tasks like medical coding, documentation and managing insurance forms cut into the time doctors can spend with patients.
Sully.ai helps solve this problem by developing AI employees that can handle routine tasks like medical coding and note-taking. As the company's platform scaled, its proprietary, closed source models created three bottlenecks: unpredictable latency in real-time clinical workflows, inference costs that scaled faster than revenue and insufficient control over model quality and updates.
Sully.ai builds AI employees that handle routine tasks for physicians. To overcome these bottlenecks, Sully.ai uses Baseten's Model API, which deploys open source models such as gpt-oss-120b on NVIDIA Blackwell GPUs. Baseten used the low-precision NVFP4 data format, the NVIDIA TensorRT-LLM library and the NVIDIA Dynamo inference framework to deliver optimized inference. The company chose NVIDIA Blackwell to run its Model API after seeing up to 2.5x better throughput per dollar compared with the NVIDIA Hopper platform.
As a result, Sully.ai's inference costs dropped by 90%, representing a 10x reduction compared with the prior closed source implementation, while response times improved by 65% for critical workflows like generating medical notes. The company has now returned over 30 million minutes to physicians, time previously lost to data entry and other manual tasks.
Gaming - DeepInfra and Latitude Reduce Cost per Token by 4x Latitude is building the future of AI-native gaming with its AI Dungeon adventure-story game and upcoming AI-powered role-playing gaming platform, Voyage, where players can create or play worlds with the freedom to choose any action and make their own story.
The company's platform uses large language models to respond to players' actions - but this comes with scaling challenges, as every player action triggers an inference request. Costs scale with engagement, and response times must stay fast enough to keep the experience seamless.
Latitude has built a text-based adventure-story game called AI Dungeon, which generates both narrative text and imagery in real time as players explore dynamic stories. Latitude runs large open source models on DeepInfra's inference platform, powered by NVIDIA Blackwell GPUs and TensorRT-LLM. For a large-scale mixture-of-experts (MoE) model, DeepInfra reduced the cost per million tokens from 20 cents on the NVIDIA Hopper platform to 10 cents on Blackwell. Moving to Blackwell's native low-precision NVFP4 format further cut that cost to just 5 cents - for a total 4x improvement in cost per token - while maintaining the accuracy that customers expect.
Running these large-scale MoE models on DeepInfra's Blackwell-powered platform allows Latitude to deliver fast, reliable responses cost effectively. DeepInfra inference platform delivers this performance while reliably handling traffic spikes, letting Latitude deploy more capable models without compromising player experience.
Agentic Chat - Fireworks AI and Sentient Foundation Lower AI Costs by up to 50% Sentient Labs is focused on bringing AI developers together to build powerful reasoning AI systems that are all open source. The goal is to accelerate AI toward solving harder reasoning problems through research in secure autonomy, agentic architecture and continual learning.
Its first app, Sentient Chat, orchestrates complex multi-agent workflows and integrates more than a dozen specialized AI agents from the community. Due to this, Sentient Chat has massive compute demands because a single user query could trigger a cascade of autonomous interactions that typically lead to costly infrastructure overhead.
To manage this scale and complexity, Sentient uses Fireworks AI's inference platform running on NVIDIA Blackwell. With Fireworks' Blackwell-optimized inference stack, Sentient achieved 25-50% better cost efficiency compared with its previous Hopper-based deployment.
Sentient Chat orchestrates complex multi-agent work
North America Stories
04/04/2026
The University of Arizona's Men's Basketball team has only loss twice th...
04/04/2026
1080p HDR arrives, a new generation of storytelling tools takes center stage, an...
04/04/2026
Michigan legends bring a new voice to the broadcast as TNT Sports and CBS Sports...
04/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
04/04/2026
DHD Introduces AI-Based Audio Noise Reduction to XD3 IP Core
Brie Clayton April 3, 2026
0 Comments
The accompanying image shows the rear panel of the ...
04/04/2026
Macnica Redefines ST 2110 Flexibility with Two Speeds on One Card
Brie Clayton April 3, 2026
0 Comments
New for NAB Show 2026, MEP100 SmartNIC now sup...
04/04/2026
Unified Media Workflows for Story-Centric Production
Brie Clayton April 3, 2026
0 Comments
Framelight X unifies field capture, editing and publishing ...
03/04/2026
Michigan's Fab Five will reunite for an alternate presentation of the Mich...
03/04/2026
Avid will exhibit at NAB Show 2026 (April 18-22, Booth N2226, Las Vegas Convention Center), demonstrating its Content Core platform and new AI-driven workflow c...
03/04/2026
Mark Roberts Motion Control (MRMC) has announced the appointment of Nick Barthee as Chief Operating Officer.
The announcement follows MRMC's transition fro...
03/04/2026
Interra Systems has announced that Elite Media Technologies has selected its BATON file-based QC solution for media workflows. Elite Media Technologies speciali...
03/04/2026
Ateme has announced that Moldtelecom has deployed Ateme technologies across its streaming workflow, covering encoding, delivery, operations, and analytics.
Mol...
03/04/2026
Grass Valley will demonstrate Framelight X, its content management platform, at NAB Show 2026. The platform connects capture, ingest, editing, and publishing in...
03/04/2026
Encompass Digital Media and Techex have announced a cloud-native Master Control ...
03/04/2026
Live Vertical Video automatically track the action on the court via AI technology and delivers a fully optimized, 9 16 live feed for viewers...
03/04/2026
As the Illini make their first trip to college basketball's biggest stage si...
03/04/2026
After last summer's Softball National Championship victory and last week'...
03/04/2026
The University of Arizona's Men's Basketball team has only loss twice th...
03/04/2026
Eight games across four tournaments will be played in three venues; accommodatio...
03/04/2026
The Ottawa Senators and Bell Media have announced a long-term rights extension for regional Ottawa Senators games on TSN and RDS. TSN Radio 1200 remains the exc...
03/04/2026
Massive production in Phoenix is run out of Game Creek Video Flagship mobile uni...
03/04/2026
New York April 2, 2026 TelevisaUnivision, the world's leading Spanish-la...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
CVP, one of Europe's leading suppliers of professional video and broadcast solutions, today announces the launch of its new German operation and the formati...
03/04/2026
Mark Roberts Motion Control (MRMC) today announces the appointment of Nick Barthee as Chief Operating Officer, strengthening its leadership as the company conti...
03/04/2026
Net Insight introduces programmable Trust Boundaries that make live media interconnection predictable as traffic moves between facilities, networks and cloud en...
03/04/2026
Winning in the new media economy: Avid showcases AI-powered, connected intellige...
03/04/2026
NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibilit...
03/04/2026
NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances
Brie Clayton April 2, 2026
0 Comments
PlayBox Neo will showcase the latest i...
03/04/2026
For Taku Hirano, Everything Is Connected From touring and composition to teaching and instrument design, the in-demand percussionist sees it all as one body o...
03/04/2026
Berklee Honors Humberto Ramirez with Master of Latin Music Award The alumnus and acclaimed trumpeter is honored for his influence as a performer, composer, an...
03/04/2026
Back to All News
Competition Heats Up with Intrigue and Spices: Netflix Unveils...
03/04/2026
Back to All News
Radioactive Emergency Ranks #1 On Netflix's Global Top 10 ...
02/04/2026
HBO and NFL Films have announced Hard Knocks: Training Camp with the Seattle Sea...
02/04/2026
Haivision has announced the Makito ONE, a single-blade video encoding and decoding platform, at NAB Show 2026. The platform combines dual-channel video encoding...
02/04/2026
Telestream has introduced UP.Lens, a cloud-based multiviewer and monitoring serv...
02/04/2026
Mark Roberts Motion Control (MRMC) will exhibit at NAB Show 2026 (Booth C5220, April 19-22, Las Vegas Convention Center), marking the company's 60th anniver...
02/04/2026
Net Insight has introduced programmable Trust Boundaries, a feature integrated i...
02/04/2026
Bitmovin has announced support for SGAI (Server-Guided Ad Insertion) in its playback products, using HLS interstitials. SGAI combines elements of client-side an...
02/04/2026
Riedel Communications' SimplyLive RiMotion R12 replay system is supporting B...
02/04/2026
LTN, a managed IP video transport company, and Ateme, a video compression and de...
02/04/2026
Harmonic has announced that TDF, a broadcast infrastructure operator in France, has deployed Harmonic's XOS Advanced Media Processor and ProStream X Video S...