Sony Pixel Power calrec Sony

Think SMART: How to Optimize AI Factory Inference Performance

21/08/2025

From AI assistants doing deep research to autonomous vehicles making split-second navigation decisions, AI adoption is exploding across industries.

Behind every one of those interactions is inference - the stage after training where an AI model processes inputs and produces outputs in real time.

Today's most advanced AI reasoning models - capable of multistep logic and complex decision-making - generate far more tokens per interaction than older models, driving a surge in token usage and the need for infrastructure that can manufacture intelligence at scale.

AI factories are one way of meeting these growing needs.

But running inference at such a large scale isn't just about throwing more compute at the problem.

To deploy AI with maximum efficiency, inference must be evaluated based on the Think SMART framework:

Scale and complexity

Multidimensional performance

Architecture and software

Return on investment driven by performance

Technology ecosystem and install base

Scale and Complexity As models evolve from compact applications to massive, multi-expert systems, inference must keep pace with increasingly diverse workloads - from answering quick, single-shot queries to multistep reasoning involving millions of tokens.

The expanding size and intricacy of AI models introduce major implications for inference, such as resource intensity, latency and throughput, energy and costs, as well as diversity of use cases.

To meet this complexity, AI service providers and enterprises are scaling up their infrastructure, with new AI factories coming online from partners like CoreWeave, Dell Technologies, Google Cloud and Nebius.

Multidimensional Performance Scaling complex AI deployments means AI factories need the flexibility to serve tokens across a wide spectrum of use cases while balancing accuracy, latency and costs.

Some workloads, such as real-time speech-to-text translation, demand ultralow latency and a large number of tokens per user, straining computational resources for maximum responsiveness. Others are latency-insensitive and geared for sheer throughput, such as generating answers to dozens of complex questions simultaneously.

But most popular real-time scenarios operate somewhere in the middle: requiring quick responses to keep users happy and high throughput to simultaneously serve up to millions of users - all while minimizing cost per token.

For example, the NVIDIA inference platform is built to balance both latency and throughput, powering inference benchmarks on models like gpt-oss, DeepSeek-R1 and Llama 3.1.

What to Assess to Achieve Optimal Multidimensional Performance

Throughput: How many tokens can the system process per second? The more, the better for scaling workloads and revenue.

Latency: How quickly does the system respond to each individual prompt? Lower latency means a better experience for users - crucial for interactive applications.

Scalability: Can the system setup quickly adapt as demand increases, going from one to thousands of GPUs without complex restructuring or wasted resources?

Cost Efficiency: Is performance per dollar high, and are those gains sustainable as system demands grow?

Architecture and Software AI inference performance needs to be engineered from the ground up. It comes from hardware and software working in sync - GPUs, networking and code tuned to avoid bottlenecks and make the most of every cycle.

Powerful architecture without smart orchestration wastes potential; great software without fast, low-latency hardware means sluggish performance. The key is architecting a system so that it can quickly, efficiently and flexibly turn prompts into useful answers.

Enterprises can use NVIDIA infrastructure to build a system that delivers optimal performance.

Architecture Optimized for Inference at AI Factory Scale The NVIDIA Blackwell platform unlocks a 50x boost in AI factory productivity for inference - meaning enterprises can optimize throughput and interactive responsiveness, even when running the most complex models.

The NVIDIA GB200 NVL72 rack-scale system connects 36 NVIDIA Grace CPUs and 72 Blackwell GPUs with NVIDIA NVLink interconnect, delivering 40x higher revenue potential, 30x higher throughput, 25x more energy efficiency and 300x more water efficiency for demanding AI reasoning workloads.

Further, NVFP4 is a low-precision format that delivers peak performance on NVIDIA Blackwell and slashes energy, memory and bandwidth demands without skipping a beat on accuracy, so users can deliver more queries per watt and lower costs per token.

Full-Stack Inference Platform Accelerated on Blackwell Enabling inference at AI factory scale requires more than accelerated architecture. It requires a full-stack platform with multiple layers of solutions and tools that can work in concert together.

Modern AI deployments require dynamic autoscaling from one to thousands of GPUs. The NVIDIA Dynamo platform steers distributed inference to dynamically assign GPUs and optimize data flows, delivering up to 4x more performance without cost increases. New cloud integrations further improve scalability and ease of deployment.

For inference workloads focused on getting optimal performance per GPU, such as speeding up large mixture of expert models, frameworks like NVIDIA TensorRT-LLM are helping developers achieve breakthrough performance.

With its new PyTorch-centric workflow, TensorRT-LLM streamlines AI deployment by removing the need for manual engine management. These solutions aren't just powerful on their own - they're built to work in tandem. For example, using Dynamo and TensorRT-LLM, mission-critical inference providers like Baseten can immediately deliver state-of-the-art model performance even on new frontier models like gpt-oss.

On the model side, families like NVIDIA Nemotron are built with open training data for t
LINK: https://blogs.nvidia.com/blog/think-smart-optimize-ai-factory-inferenc...
See more stories from nvidia

Most recent headlines

06/10/2025

France Tlvisions Wins Prestigious 2025 EBU Technology & Innovation Award in Groundbreaking Collaboration with Dalet

France T l visions, France's leading broadcaster, has received the 2025 EBU ...

04/09/2025

Monumental Sports & Entertainment and Dalet Win Prestigious 2025 NAB Show Project of the Year Award

Monumental Sports & Entertainment (MSE), in collaboration with Dalet, has been a...

21/08/2025

Clear-Com Connects the Desert Across Coachella and Stagecoach 2025

eds3_5_jq(document).ready(function($) { $(#eds_sliderM519).chameleonSlider_2_1({ content_source:......

21/08/2025

Imagine Communications to Debut SNP-XS at IBC2025

DENVER At IBC2025, Sept. 12-15 at the RAI Amsterdam, Imagine Communications will introduce the SNP-XS, a versatile addition to its Selenio Network Processor (SN...

21/08/2025

BitFire Launches Live Master Control in the Cloud

HUDSON, Mass. BitFire, a provider of software-defined live production and IP transmission, today announced the addition of cloud-based live master control capab...

21/08/2025

Gray Media to Launch New Hyper-Personalized Video Streaming Service

ATLANTA Gray Media has laid out plans for launching a new cutting-edge, hyper-personalized streaming platform that will start going live in Grays markets in Jan...

21/08/2025

Chyron Partners with Asport for Live Sports Production and Distribution

NEW YORK and ZURICH Chyron, a provider of broadcast graphics and live production solutions, has announced a partnership with Asport, a leading sports tech innov...

21/08/2025

swXtch.io to Debut SRTx Gateway at IBC2025

NEW YORK swXtch.io has amplified its support for SRT workflows with a new specialized gateway solution primarily targeted at the live event market. To be introd...

21/08/2025

Rise Academy Revives 4K Charity Run for IBC 2025

AMSTERDAM Rise Academy, the charity dedicated to delivering practical media technology experiences, careers resources and sharing work experience opportunities ...

21/08/2025

Carr Names a New Special Assistant

WASHINGTON FCC Chairman Brendan Carr announced the appointment of Courtney Cowper as a special assistant in his office. As a special assistant in the Office of ...

21/08/2025

Live Media Group Debuts New IP-based REMI Production Truck

COLUMBUS, Ohio Live Media Group has launched its latest mobile production unit, the MU-28, which is a SMPTE 2110-7 IP-based truck built specifically for remote ...

21/08/2025

Sling TV Launches Sling Select

ENGLEWOOD, Colo. Sling TV has launched a new offering called Select that provides a package of cable channels for $19.99 a month....

21/08/2025

Viant and Wurl Partner on Scene-Level CTV Targeting and Measurement

IRVINE, Calif. CTV and programmatic ad provider Viant Technology Inc. has announced a new integration of its DSP with Wurl that provides advertisers with scene...

21/08/2025

Marshall Electronics to Show New PTZ Camera at IBC2025

TORRANCE, Calif. Marshall Electronics will highlight several new products at IBC2025, including the CV612 PTZ camera, RCP Plus camera controller and VMV-402-3GS...

21/08/2025

Think SMART: How to Optimize AI Factory Inference Performance

From AI assistants doing deep research to autonomous vehicles making split-second navigation decisions, AI adoption is exploding across industries. Behind ever...

21/08/2025

Gearing Up for the Gigawatt Data Center Age

Across the globe, AI factories are rising - massive new data centers built not to serve up web pages or email, but to train and deploy intelligence itself. Inte...

21/08/2025

COMING SOON TO RT More, More, More with Your Public Service Media

RT today announces its exciting upcoming slate of video content, taking viewers from Autumn 2025 to Spring 2026 across RT One, RT 2 and RT Player. From an un...

21/08/2025

GeForce NOW Brings RTX 5080 Power to the Ultimate Membership

Get a glimpse into the future of gaming. The NVIDIA Blackwell RTX architecture is coming to GeForce NOW in September, marking the service's biggest upgrade...

21/08/2025

Mahindra Sets a New Benchmark: XUV 3XO REVX A Becomes the World's First SUV Under 12 Lakh to Feature Dolby Atmos

August 21 2025, 00:13 (PDT) Mahindra Sets a New Benchmark: XUV 3XO REVX A Becom...

20/08/2025

8 Must-Listen Podcast Episodes to Fuel This Summer's Blockbuster Hype

There's still plenty of summer left to soak up-and time to catch its hottest flicks. This year, blockbuster season is alive with remakes, superhero sagas, a...

20/08/2025

SBS unleashes landmark docu-drama The People vs Robodebt

SBS unleashes landmark docu-drama The People vs Robodebt 20 August, 2025 Media releases Our Government. Half a Million Victims [1]. Zero Accountability Un...

20/08/2025

Dan Bourchier appointed General Manager of NITV

Dan Bourchier appointed General Manager of NITV 20 August, 2025 Media releases The respected journalist, presenter and accomplished leader will join Nation...

20/08/2025

Building Hollywood's Village: HPA President Kari Grubin on Community, Innovation, and Change

By Daron James. Originally published, August 13, 2025 on motionpictures.org. Si...

20/08/2025

US Military Uses L3Harris Electronic Warfare System During Bilateral Exercise

L3harris successfully demonstrates DiSCO at Talisman Sabre....

20/08/2025

The Enduring Legacy of L3Harris Propulsion on Viking Missions

Viking 1 was the United States' first attempt to soft-land a spacecraft on another planet and it provided the first photograph ever taken on the surface of ...

20/08/2025

Calrec Meets NEP UK's Goals for EFL Coverage

Simplifying Sky Sports' EFL production coverage, NEP UK's NEO hybrid outside broadcast trailer relies on Calrec remote workflows to minimise production ...

20/08/2025

Sky and TVNZ back Nielsen as NZ TV Measurement provider under new partnership

Auckland, New Zealand - 19 August 2025: Nielsen, a global leader in audience measurement, data, and analytics, is pleased to announce a two-year extension for T...

20/08/2025

The Gauge: Poland | July 2025

July, much like June, is a month marked by summer weather and vacations, which has reflected in a further decline in viewer activity in front of TV screens. The...

20/08/2025

THE GAUGE: MEXICO JULY 2025

During July, streaming's share of TV viewing in Mexico showed an increase of 0.9 percentage points compared to the previous month, accounting for 24.6% of T...

20/08/2025

Streaming Cranks Up the Heat in July, Accounts For Nearly Half of All TV Viewing in Nielsen's The Gauge

Netflix, YouTube and The Roku Channel Each Hit Platform Highs Netflix Owns 8 of...

20/08/2025

InfoSum Integrates Nielsen Marketing Cloud Data For Audience Engagement

New collaboration brings granular audience insights from NMC so advertisers, agencies, and publishers can gain a better understanding of their audience Now ava...

20/08/2025

ASG Opens New Burbank Facility

BURBANK, Calif. Advanced Systems Group has opened a 5,000 square-foot-facility in Burbank with a Dolby Atmos demo theater, a video production space with a hard ...

20/08/2025

BHV Brings First Production Units of SportsBox to IBC 202...

BHV, renowned for its design and engineering expertise behind many of broadcast's leading names, returns to IBC with the first production units of SportsBox...

20/08/2025

Marshall Electronics to Display Range of New Products at...

Marshall Electronics will highlight several of its new product offerings at IBC 2025 (Booth 11.C28), including the CV612 PTZ Camera, RCP Plus Camera Controller ...

20/08/2025

LTN accelerates satellite-to-IP transition with new cost...

LTN announces a range of enhancements to its purpose-built IP video network as the broadcast industry approaches a major inflection point in the shift from sate...

20/08/2025

Yahoo Sports to Launch New Streaming FAST Channel

NEW YORK Yahoo Sports and C15 Studio have announced the launch of the Yahoo Sports Network, a free, ad-supported streaming TV channel available on leading FAST ...

20/08/2025

Gravity Media expands Calrec technology at London Product...

Completing its latest expansion phase, Gravity Media's Production Centre facility in London White City has extended its remote and distributed production ca...

20/08/2025

HighField AI Expands Global Team Following Commercial Ava...

HighField AI, the broadcast industry's first agentic and multimodal AI platform for automated graphics production, today announced the expansion of its comm...

20/08/2025

Zixi to Showcase Scalable IP Video Infrastructure for Liv...

Zixi, the Emmy Award-winning leader in broadcast-quality live video over IP, will exhibit at IBC 2025 in Amsterdam (September 12 15) at Stand 5.A85. The compan...

20/08/2025

Rise Academy Revives Race 4 the Future at IBC to Champion...

Rise Academy, the charity dedicated to delivering practical media technology experiences, careers resources and sharing work experience opportunities for young ...

20/08/2025

JFM Streams Over 4000 Local Sports Matches a Year Revolut...

Jysk Fynske Medier (JFM), Denmark's second-largest media group, Jysk Fynske Medier (JFM), has teamed up with Wowza's Flowplayer to deliver over 4,000 li...

20/08/2025

Streaming Heats Up in July, Accounts For Nearly Half of All TV Viewing

NEW YORK Streaming viewership continued to heat up this summer in July as its portion of the time spent watching TV edged closer to the 50% threshold, according...

20/08/2025

EVS to Acquire Telemetrics

SERAING, Belgium EVS has announced that it is acquiring U.S.-based Telemetrics Inc., a pioneer in robotics for media production....

20/08/2025

Videndum Readies For IBC2025 With Products From Multiple Brands

BURY ST. EDMUNDS, U.K. Videndum will highlight the latest developments from its broad range of product brands, such as Anton/Bauer's EDEN clean-energy alter...

20/08/2025

Pliant Technologies to Showcase Latest Intercom Solutions at IBC2025

AMSTERDAM Pliant Technologies has announced that it will be a range of its latest intercom technologies at a IBC 2025 (Booth 10.F33), including two new addition...

20/08/2025

Wooden Camera Releases Accessory Collection for Sony FX2

Wooden Camera announces the release of its new Accessory Collection for the Sony FX2. Designed with modularity and affordability in mind, the collection introdu...

20/08/2025

Triveni Digital to Showcase Comprehensive TV 3 0 Lineup a...

Triveni Digital, a trusted leader in ATSC 3.0 service delivery, data broadcasting, and quality assurance solutions, today announced it is showcasing a comprehen...

20/08/2025

Jules Woods Streamlines High-End TV and Film Sound Workfl...

With a career that began in the backrooms of post-production facilities and evolved into leading sound for major scripted series and films, Jules Woods has made...

20/08/2025

Friend MTS awarded 2025 Best Practices Technology Innovat...

Friend MTS, the number one anti-piracy provider and video cybersecurity partner in entertainment, media and sports, today announced that it has been presented w...

20/08/2025

U&Dave announces Hit Point, the channel's first ever original drama series

U&Dave today announces its first ever original drama series, Hit Point (6x60'), which will also be available to stream on U. Written by BAFTA-winner Howard ...