
A diagnostic insight in healthcare. A character's dialogue in an interactive game. An autonomous resolution from a customer service agent. Each of these AI-powered interactions is built on the same unit of intelligence: a token.
Scaling these AI interactions requires businesses to consider whether they can afford more tokens. The answer lies in better tokenomics - which at its core is about driving down the cost of each token. This downward trend is unfolding across industries. Recent MIT research found that infrastructure and algorithmic efficiencies are reducing inference costs for frontier-level performance by up to 10x annually.
To understand how infrastructure efficiency improves tokenomics, consider the analogy of a high-speed printing press. If the press produces 10x output with incremental investment in ink, energy and the machine itself, the cost to print each individual page drops. In the same way, investments in AI infrastructure can lead to far greater token output compared with the increase in cost - causing a meaningful reduction in the cost per token.
When token output outpaces infrastructure cost, the cost of each token drops. That's why leading inference providers including Baseten, DeepInfra, Fireworks AI and Together AI are using the NVIDIA Blackwell platform, which helps them reduce cost per token by up to 10x compared with the NVIDIA Hopper platform.
These providers host advanced open source models, which have now reached frontier-level intelligence. By combining open source frontier intelligence, the extreme hardware-software codesign of NVIDIA Blackwell and their own optimized inference stacks, these providers are enabling dramatic token cost reductions for businesses across every industry.
Healthcare - Baseten and Sully.ai Cut AI Inference Costs by 10x In healthcare, tedious, time-consuming tasks like medical coding, documentation and managing insurance forms cut into the time doctors can spend with patients.
Sully.ai helps solve this problem by developing AI employees that can handle routine tasks like medical coding and note-taking. As the company's platform scaled, its proprietary, closed source models created three bottlenecks: unpredictable latency in real-time clinical workflows, inference costs that scaled faster than revenue and insufficient control over model quality and updates.
Sully.ai builds AI employees that handle routine tasks for physicians. To overcome these bottlenecks, Sully.ai uses Baseten's Model API, which deploys open source models such as gpt-oss-120b on NVIDIA Blackwell GPUs. Baseten used the low-precision NVFP4 data format, the NVIDIA TensorRT-LLM library and the NVIDIA Dynamo inference framework to deliver optimized inference. The company chose NVIDIA Blackwell to run its Model API after seeing up to 2.5x better throughput per dollar compared with the NVIDIA Hopper platform.
As a result, Sully.ai's inference costs dropped by 90%, representing a 10x reduction compared with the prior closed source implementation, while response times improved by 65% for critical workflows like generating medical notes. The company has now returned over 30 million minutes to physicians, time previously lost to data entry and other manual tasks.
Gaming - DeepInfra and Latitude Reduce Cost per Token by 4x Latitude is building the future of AI-native gaming with its AI Dungeon adventure-story game and upcoming AI-powered role-playing gaming platform, Voyage, where players can create or play worlds with the freedom to choose any action and make their own story.
The company's platform uses large language models to respond to players' actions - but this comes with scaling challenges, as every player action triggers an inference request. Costs scale with engagement, and response times must stay fast enough to keep the experience seamless.
Latitude has built a text-based adventure-story game called AI Dungeon, which generates both narrative text and imagery in real time as players explore dynamic stories. Latitude runs large open source models on DeepInfra's inference platform, powered by NVIDIA Blackwell GPUs and TensorRT-LLM. For a large-scale mixture-of-experts (MoE) model, DeepInfra reduced the cost per million tokens from 20 cents on the NVIDIA Hopper platform to 10 cents on Blackwell. Moving to Blackwell's native low-precision NVFP4 format further cut that cost to just 5 cents - for a total 4x improvement in cost per token - while maintaining the accuracy that customers expect.
Running these large-scale MoE models on DeepInfra's Blackwell-powered platform allows Latitude to deliver fast, reliable responses cost effectively. DeepInfra inference platform delivers this performance while reliably handling traffic spikes, letting Latitude deploy more capable models without compromising player experience.
Agentic Chat - Fireworks AI and Sentient Foundation Lower AI Costs by up to 50% Sentient Labs is focused on bringing AI developers together to build powerful reasoning AI systems that are all open source. The goal is to accelerate AI toward solving harder reasoning problems through research in secure autonomy, agentic architecture and continual learning.
Its first app, Sentient Chat, orchestrates complex multi-agent workflows and integrates more than a dozen specialized AI agents from the community. Due to this, Sentient Chat has massive compute demands because a single user query could trigger a cascade of autonomous interactions that typically lead to costly infrastructure overhead.
To manage this scale and complexity, Sentient uses Fireworks AI's inference platform running on NVIDIA Blackwell. With Fireworks' Blackwell-optimized inference stack, Sentient achieved 25-50% better cost efficiency compared with its previous Hopper-based deployment.
Sentient Chat orchestrates complex multi-agent work
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
18/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
18/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
18/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
18/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
18/06/2026
Linda May Han Oh Receives Guggenheim Fellowship The Berklee professor, bassist, and composer will use the fellowship to debut Dreams of Knowing, an interdisci...
17/06/2026
EVS has announced it has received the EcoVadis Gold Medal for sustainability performance, ranking among the top 5% of companies globally in the Technology/Mid-S...
17/06/2026
Chyron has released Weather 2.4, an update to its weather suite for broadcasters and meteorologists. The release focuses on enhancements to the DataFlow module,...
17/06/2026
VSiN, The Sports Betting Network, has announced the launch of Best Bets TV, a free ad-supported streaming TV (FAST) channel. The 24/7 channel is currently avail...
17/06/2026
DAZN's Team Whistle and Snap Inc. have announced a creator program centered ...
17/06/2026
LiveU is providing video transmission technology for broadcasters, production companies, and public safety agencies across North America's busy Summer of So...
17/06/2026
The National Academy of Television Arts and Sciences (NATAS) has announced that Laurens Grant and Jacob Ullman have joined its Board of Directors. Chief of Staf...
17/06/2026
Akta, the AI-First SaaS video platform for modern broadcast and streaming operations, today announced that its video platform is now generally available on Orac...
17/06/2026
This recent graduate from Houston found inspiration in technical directing and now eyes a future career in sports production...
17/06/2026
Audio-Technica (booth C7959) arrives at InfoComm 2026 in Las Vegas with a slate ...
17/06/2026
SNS has published a guide addressing growing demand for AI-powered video indexing, transcription, facial recognition, and searchable metadata across media libra...
17/06/2026
Providius has announced Providius Direct, a workflow for investigating network i...
17/06/2026
NEP Group has announced the commercial availability of NEP Platform, a software ...
17/06/2026
A new white paper examining Secure Reliable Transport (SRT) and Reliable Interne...
17/06/2026
New research from subscription bundling platform Bango finds that younger sports fans are increasingly consuming sport through highlights, clips, and social med...
17/06/2026
SMPTE has announced that its complete Standards catalog is now freely available to the global media technology community, including all published SMPTE Standard...
17/06/2026
Harmonic has completed the sale of its Video Business to MediaKind for $145 mill...
17/06/2026
Omaha Productions will produce the 2026 World Series of Poker (WSOP) in Las Vega...
17/06/2026
New features, changes & bug fixes
SoundBridge have just released another update for their remote collaboration-focused DAW - reviewed here in SOS March 2026...
17/06/2026
Valve-based front end for digital & modelling rigs
The latest addition to Fryette's product range delivers a packed-down, pedalboard-friendly version of...
17/06/2026
GearExpo UK - 27 June 2026
Sound On Sound are proud to announce GearExpo UK, a major new recording and music technology exhibition in London! This is the bi...
17/06/2026
SAM monitoring line-up gains Dante and AES67 support
The latest expansion of Genelec's UNIO monitoring ecosystem introduces a new device that provides D...
17/06/2026
The R&S PR300 portable receiver from Rohde & Schwarz sets new standards in spect...
17/06/2026
Elt Group and Rohde & Schwarz sign a cooperation agreement to explore commercial...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
The Immersive Supervisor Emerges as Hollywood's Next Production Role
Brie Clayton June 17, 2026
0 Comments
Above image: On set live immersive revi...
17/06/2026
Vertical Musical Playback Shot with Blackmagic PYXIS 6K
Brie Clayton June 17, 2026
0 Comments
Large format sensor and DaVinci Resolve workflow used fo...
17/06/2026
DAZ 3D Launches New Game-Ready Character Assets Built for Modern Engines and Pro...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/06/2026
Montr al, Quebec, June 11, 2026 Kiloview, a leading provider of AV-over-IP and NDI -based video transmission solutions, today announced a distribution partner...
17/06/2026
Changsha, China, June 15, 2026 Kiloview officially announced the launch of U4 IP Video Dock, a compact IP video decoder and output dock designed to bring prof...
17/06/2026
Save 50% on All Ivory II Pianos and Collections
Were halfway through the year, and we're cutting prices in half for the Ivory II Summer Sale!
For a limite...
17/06/2026
Wednesday 17 June 2026
Two in three fans will connect to venue WiFi this World ...
17/06/2026
Visibility builds credibility - the tools you use every day, now visible on your LinkedIn profile Published on Jun 17, 2026 Categories: Company News, Product ...
17/06/2026
Transaction Positions Harmonic as a Pure-Play Broadband Company SAN JOSE, Calif. - June 17, 2026 - Harmonic Inc. (NASDAQ: HLIT), the worldwide leader in virtual...
17/06/2026
FOX Advertising To Launch Industry's First End-to-End Agentic Advertising Pl...
17/06/2026
How SGN is future-proofing critical national infrastructure with Arqiva Managed Connectivity.
When disruption becomes the norm As Storm Eunice tore across the ...