
A diagnostic insight in healthcare. A character's dialogue in an interactive game. An autonomous resolution from a customer service agent. Each of these AI-powered interactions is built on the same unit of intelligence: a token.
Scaling these AI interactions requires businesses to consider whether they can afford more tokens. The answer lies in better tokenomics - which at its core is about driving down the cost of each token. This downward trend is unfolding across industries. Recent MIT research found that infrastructure and algorithmic efficiencies are reducing inference costs for frontier-level performance by up to 10x annually.
To understand how infrastructure efficiency improves tokenomics, consider the analogy of a high-speed printing press. If the press produces 10x output with incremental investment in ink, energy and the machine itself, the cost to print each individual page drops. In the same way, investments in AI infrastructure can lead to far greater token output compared with the increase in cost - causing a meaningful reduction in the cost per token.
When token output outpaces infrastructure cost, the cost of each token drops. That's why leading inference providers including Baseten, DeepInfra, Fireworks AI and Together AI are using the NVIDIA Blackwell platform, which helps them reduce cost per token by up to 10x compared with the NVIDIA Hopper platform.
These providers host advanced open source models, which have now reached frontier-level intelligence. By combining open source frontier intelligence, the extreme hardware-software codesign of NVIDIA Blackwell and their own optimized inference stacks, these providers are enabling dramatic token cost reductions for businesses across every industry.
Healthcare - Baseten and Sully.ai Cut AI Inference Costs by 10x In healthcare, tedious, time-consuming tasks like medical coding, documentation and managing insurance forms cut into the time doctors can spend with patients.
Sully.ai helps solve this problem by developing AI employees that can handle routine tasks like medical coding and note-taking. As the company's platform scaled, its proprietary, closed source models created three bottlenecks: unpredictable latency in real-time clinical workflows, inference costs that scaled faster than revenue and insufficient control over model quality and updates.
Sully.ai builds AI employees that handle routine tasks for physicians. To overcome these bottlenecks, Sully.ai uses Baseten's Model API, which deploys open source models such as gpt-oss-120b on NVIDIA Blackwell GPUs. Baseten used the low-precision NVFP4 data format, the NVIDIA TensorRT-LLM library and the NVIDIA Dynamo inference framework to deliver optimized inference. The company chose NVIDIA Blackwell to run its Model API after seeing up to 2.5x better throughput per dollar compared with the NVIDIA Hopper platform.
As a result, Sully.ai's inference costs dropped by 90%, representing a 10x reduction compared with the prior closed source implementation, while response times improved by 65% for critical workflows like generating medical notes. The company has now returned over 30 million minutes to physicians, time previously lost to data entry and other manual tasks.
Gaming - DeepInfra and Latitude Reduce Cost per Token by 4x Latitude is building the future of AI-native gaming with its AI Dungeon adventure-story game and upcoming AI-powered role-playing gaming platform, Voyage, where players can create or play worlds with the freedom to choose any action and make their own story.
The company's platform uses large language models to respond to players' actions - but this comes with scaling challenges, as every player action triggers an inference request. Costs scale with engagement, and response times must stay fast enough to keep the experience seamless.
Latitude has built a text-based adventure-story game called AI Dungeon, which generates both narrative text and imagery in real time as players explore dynamic stories. Latitude runs large open source models on DeepInfra's inference platform, powered by NVIDIA Blackwell GPUs and TensorRT-LLM. For a large-scale mixture-of-experts (MoE) model, DeepInfra reduced the cost per million tokens from 20 cents on the NVIDIA Hopper platform to 10 cents on Blackwell. Moving to Blackwell's native low-precision NVFP4 format further cut that cost to just 5 cents - for a total 4x improvement in cost per token - while maintaining the accuracy that customers expect.
Running these large-scale MoE models on DeepInfra's Blackwell-powered platform allows Latitude to deliver fast, reliable responses cost effectively. DeepInfra inference platform delivers this performance while reliably handling traffic spikes, letting Latitude deploy more capable models without compromising player experience.
Agentic Chat - Fireworks AI and Sentient Foundation Lower AI Costs by up to 50% Sentient Labs is focused on bringing AI developers together to build powerful reasoning AI systems that are all open source. The goal is to accelerate AI toward solving harder reasoning problems through research in secure autonomy, agentic architecture and continual learning.
Its first app, Sentient Chat, orchestrates complex multi-agent workflows and integrates more than a dozen specialized AI agents from the community. Due to this, Sentient Chat has massive compute demands because a single user query could trigger a cascade of autonomous interactions that typically lead to costly infrastructure overhead.
To manage this scale and complexity, Sentient uses Fireworks AI's inference platform running on NVIDIA Blackwell. With Fireworks' Blackwell-optimized inference stack, Sentient achieved 25-50% better cost efficiency compared with its previous Hopper-based deployment.
Sentient Chat orchestrates complex multi-agent work
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
01/04/2026
January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION
Douyin Users Can Now Create And Share Videos With Stun...
29/03/2026
Cloud-based production, real-time engagement, and creator-driven storytelling ai...
28/03/2026
Now features DiGiCo console integration
Harrison's live recording and virtual soundcheck software has just reached its third major version, which among ...
28/03/2026
MPE-capable chamber strings library announced
Alongside their collection of Kontakt instruments, Sonora Cinematic have been steadily introducing a series of...
28/03/2026
Globecast, the leading provider of broadcast, media and entertainment managed services, will showcase its reimagined approach to media operations at the 2026 NA...
28/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
In-venue and creative video staffers at the professional and collegiate level ha...
27/03/2026
Comcast Business deployed network infrastructure for the 2026 PLAYERS Championsh...
27/03/2026
Czech production company CS live has equipped its newest outside broadcast van w...
27/03/2026
Edith Cowan University (ECU) in Perth, Western Australia has developed new broad...
27/03/2026
Deltatre has announced that CEO Andrea Marini will step down after five years in...
27/03/2026
DAZN has announced plans to launch DAZN Inflight, a live sports service for airline and maritime passengers, slated for 2027. Aviv Giladi, President of DAZN Par...
27/03/2026
Grass Valley has announced the completion of a live production deployment with NVP, a European media company specializing in live sports production, for LALIGA ...
27/03/2026
The Masters and Prime Video will debut Inside Amen Corner, a dedicated feed that...
27/03/2026
ESPN and the World Series of Poker (WSOP) have reached a multi-year agreement to bring the WSOP Main Event back to ESPN platforms. Coverage will include a three...
27/03/2026
USSI Global has opened its Media Transport Solutions Lab on its Melbourne campus. The engineering center provides a platform-agnostic environment for testing al...
27/03/2026
From 14-camera coverage to official review from Variant Systems Group, Sellitto ...
27/03/2026
The United Football League (UFL) has named Sportable its Official Connected Ball and Player Tracking Partner. Sportable's connected football and wearable pl...
27/03/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
27/03/2026
The Atlanta Braves and FuboTV have announced a multiyear distribution agreement to carry BravesVision on Fubo's live TV streaming platform beginning Opening...
27/03/2026
Besides restructuring for Season 3, the league worked with FOX and ESPN over the...
27/03/2026
The Atlanta Braves open their 2026 MLB season tonight against the Kansas City Ro...
27/03/2026
With new teams and new venues to adapt to, the spring-football league's part...
27/03/2026
By Lucy Spicer
One of the most exciting things about the Sundance Film Festival...
27/03/2026
New Classic, Amplitude & FlexRange units introduced
GIK Acoustics have just introduced a trio of new bass trap designs that bring improved low-end absorptio...
27/03/2026
Offers personalised Dolby Atmos headphone monitoring
Sonarworks have put their calibration expertise to work on a new mobile app that allows users to create...
27/03/2026
NITV to broadcast farewell to Rhoda Roberts AO with special coverage and week-lo...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
Marshall Electronics Showcases New Feature-Rich CV320 and CV520 IP and 3G-SDI PO...
27/03/2026
Sony Electronics Inc. Elevates Professional Video Workflows with Powerful Update...
27/03/2026
GatesAir Extends AirWatch365 Managed Service with Edge Gateway Site Appliance
Brie Clayton March 27, 2026
0 Comments
NAB marks global launch of servic...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/03/2026
At NAB Show 2026, Net Insight will showcase the next evolution of Nimbra Edge, its orchestration and control layer designed to manage live media services across...
27/03/2026
Harmonic (NASDAQ: HLIT) today announced powerful new innovations that further elevate the company's sports streaming solution. The advanced capabilities enh...
27/03/2026
Bitmovin, the leading provider of video streaming solutions, today announced significant new capabilities for Player Web X, its next-generation web video player...
27/03/2026
Riedel Communications today announced that Czech-based production company CS live has equipped its newest outside broadcast (OB) van with an integrated Riedel i...
27/03/2026
130 Industry Experts Confirmed as Show Celebrates its 10th Anniversary
MPTS, the UK's largest and most influential event for the media, production and tech...
27/03/2026
Grass Valley today announced the successful completion of a major live production deployment with NVP, a leading European media company specializing in live spo...
27/03/2026
Showcases resilient solutions for satellite-to-IP migration, REMI and hybrid live production
Appear ASA (Appear, OSE:APR), a global leader in live production t...
27/03/2026
San Francisco, California, March 2026 - Microsoft Ignite, a major annual conference hosted by Microsoft for developers, IT professionals, partners and business ...
27/03/2026
Clear-Com is proud to highlight its support for worship teams through professional communication solutions, with the deployment of its EQUIP wireless system a...