Sony Pixel Power calrec Sony

Think SMART: New NVIDIA Dynamo Integrations Simplify AI Inference at Data Center Scale

10/11/2025

Editor's note: This post is part of Think SMART, a series focused on how leading AI service providers, developers and enterprises can boost their inference performance and return on investment with the latest advancements from NVIDIA's full-stack inference platform.

AI models are becoming increasingly complex and collaborative through multi-agent workflows. To keep up, AI inference must now scale across entire clusters to serve millions of concurrent users and deliver faster responses.

Much like it did for large-scale AI training, Kubernetes - the industry standard for containerized application management - is well-positioned to manage the multi-node inference needed to support advanced models.

The NVIDIA Dynamo platform works together with Kubernetes to streamline the management of both single- and multi-node AI inference. Read on to learn how the shift to multi-node inference is driving performance, as well as how cloud platforms are putting these technologies to work.

Tapping Disaggregated Inference for Optimized Performance For AI models that fit on a single GPU or server, developers often run many identical replicas of the model in parallel across multiple nodes to deliver high throughput. In a recent paper, Russ Fellows, principal analyst at Signal65, showed that this approach achieved an industry-first record aggregate throughput of 1.1 million tokens per second with 72 NVIDIA Blackwell Ultra GPUs.

When scaling AI models to serve many concurrent users in real time, or when managing demanding workloads with long input sequences, using a technique called disaggregated serving unlocks further performance and efficiency gains.

Serving AI models involves two phases: processing the input prompt (prefill) and generating the output (decode). Traditionally, both phases run on the same GPUs, which can create inefficiencies and resource bottlenecks.

Disaggregated serving solves this by intelligently assigning these tasks to independently optimized GPUs. This approach ensures that each part of the workload runs with the optimization techniques best suited for it, maximizing overall performance. For today's large AI reasoning models, such as DeepSeek-R1, disaggregated serving is essential.

NVIDIA Dynamo seamlessly brings multi-node inference optimization features such as disaggregated serving to production scale across GPU clusters.

It's already delivering value.

Baseten, for example, used NVIDIA Dynamo to speed up inference serving for long-context code generation by 2x and increase throughput by 1.6x, all without incremental hardware costs. Such software-driven performance boosts enable AI providers to significantly reduce the costs to manufacture intelligence.

In addition, recent SemiAnalysis InferenceMAX benchmarks demonstrated that disaggregated serving with Dynamo on NVIDIA GB200 NVL72 systems delivers the lowest cost per million tokens for mixture-of-experts reasoning models like DeepSeek-R1, among platforms tested.

Scaling Disaggregated Inference in the Cloud As disaggregated serving scales across dozens or even hundreds of nodes for enterprise-scale AI deployments, Kubernetes provides the critical orchestration layer. With NVIDIA Dynamo now integrated into managed Kubernetes services from all major cloud providers, customers can scale multi-node inference across NVIDIA Blackwell systems, including GB200 and GB300 NVL72, with the performance, flexibility and reliability that enterprise AI deployments demand.

Amazon Web Services is accelerating generative AI inference for its customers with NVIDIA Dynamo and integrated with Amazon EKS.

Google Cloud is providing a NVIDIA Dynamo recipe to optimize large language model (LLM) inference at enterprise scale on its AI Hypercomputer.

OCI is enabling multi-node LLM inferencing with OCI Superclusters and NVIDIA Dynamo.

The push towards enabling large-scale, multi-node inference extends beyond hyperscalers.

Nebius, for example, is designing its cloud to serve inference workloads at scale, built on NVIDIA accelerated computing infrastructure and working with NVIDIA Dynamo as an ecosystem partner.

Simplifying Inference on Kubernetes With NVIDIA Grove in NVIDIA Dynamo Disaggregated AI inference requires coordinating a team of specialized components - prefill, decode, routing and more - each with different needs. The challenge for Kubernetes is no longer about running more parallel copies of a model, but rather about masterfully conducting these distinct components as one cohesive, high-performance system.

NVIDIA Grove, an application programming interface now available within NVIDIA Dynamo, allows users to provide a single, high-level specification that describes their entire inference system.

For example, in that single specification, a user could simply declare their requirements: I need three GPU nodes for prefill and six GPU nodes for decode, and I require all nodes for a single model replica to be placed on the same high-speed interconnect for the quickest possible response.

From that specification, Grove automatically handles all the intricate coordination: scaling related components together while maintaining correct ratios and dependencies, starting them in the right order and placing them strategically across the cluster for fast, efficient communication. Learn more about how to get started with NVIDIA Grove in this technical deep dive.

As AI inference becomes increasingly distributed, the combination of Kubernetes and NVIDIA Dynamo with NVIDIA Grove simplifies how developers build and scale intelligent applications.

Explore how these technologies come together to make cluster-scale AI easy and production-ready by joining NVIDIA at KubeCon, running through Thursday, Nov. 13, in Atlanta.
LINK: https://blogs.nvidia.com/blog/think-smart-dynamo-ai-inference-data-cen...
See more stories from nvidia

More from Nvidia

28/02/2026

NVIDIA and Partners Show That Software-Defined AI-RAN Is the Next Wireless Generation

AI-RAN is moving from lab to field, showing that a software-defined approach is ...

28/02/2026

NVIDIA Advances Autonomous Networks With Agentic AI Blueprints and Telco Reasoning Models

Autonomous networks - intelligent, self-managing telecommunications operations -...

26/02/2026

The Nightmare Returns in the Cloud: GeForce NOW Unleashes Capcom's Resident Evil Requiem'

GeForce NOW's anniversary celebration reaches a chilling crescendo as Capcom...

26/02/2026

Horror Awakens in the Cloud: GeForce NOW Unleashes Capcom's Resident Evil: Requiem'

GeForce NOW's anniversary celebration reaches a chilling crescendo as Capcom...

24/02/2026

From Radiology to Drug Discovery, Survey Reveals AI Is Delivering Clear Return on Investment in Healthcare

AI is accelerating every aspect of healthcare - from radiology and drug discover...

23/02/2026

NVIDIA Brings AI-Powered Cybersecurity to World's Critical Infrastructure

As technologies and systems become more digitalized and connected across the world, operational technology (OT) environments and industrial control systems (ICS...

19/02/2026

All About the Games: Play Over 4,500 Titles With GeForce NOW

The GeForce NOW anniversary celebration keeps on rolling, and this week is all about the games that make it possible. With more than 4,500 titles supported in t...

19/02/2026

Survey Reveals AI Advances in Telecom: Networks and Automation in Driver's Seat as Return on Investment Climbs

AI is accelerating the telecommunications industry's transformation, becomin...

17/02/2026

NVIDIA and Global Industrial Software Leaders Partner With India's Largest Manufacturers to Drive AI Boom

India is entering a new age of industrialization, as AI transforms how the world...

17/02/2026

India Fuels Its AI Mission With NVIDIA

India is the nexus of AI innovation this week as the host of the AI Impact Summit, which brings together global heads of state and industry to chart the future ...

16/02/2026

New Data Shows NVIDIA Blackwell Ultra Delivers up to 50x Better Performance and 35x Lower Costs for Agentic AI

The NVIDIA Blackwell platform has been widely adopted by leading inference provi...

12/02/2026

NVIDIA DGX Spark Powers Big Projects in Higher Education

At leading institutions across the globe, the NVIDIA DGX Spark desktop supercomputer is bringing data center class AI to lab benches, faculty offices and studen...

12/02/2026

Leading Inference Providers Cut AI Costs by up to 10x With Open Source Models on NVIDIA Blackwell

A diagnostic insight in healthcare. A character's dialogue in an interactive...

12/02/2026

GeForce NOW Turns Screens Into a Gaming Machine

The GeForce NOW sixth-anniversary festivities roll on this February, continuing a monthlong celebration of NVIDIA's cloud gaming service. This week brings ...

05/02/2026

GeForce NOW Celebrates Six Years of Streaming With 24 Games in February

Break out the cake and green sprinkles - GeForce NOW is turning six. Since launch, members have streamed over 1 billion hours, and the party's just getting...

04/02/2026

Nemotron Labs: How AI Agents Are Turning Documents Into Real-Time Business Intelligence

Editor's note: This post is part of the Nemotron Labs blog series, which exp...

03/02/2026

Everything Will Be Represented in a Virtual Twin, NVIDIA CEO Jensen Huang Says at 3DEXPERIENCE World

At 3DEXPERIENCE World in Houston, NVIDIA founder and CEO Jensen Huang and Dassau...

29/01/2026

Mercedes-Benz Unveils New S-Class Built on NVIDIA DRIVE AV, Which Enables an L4-Ready Architecture

Mercedes-Benz is marking 140 years of automotive innovation with a new S-Class b...

29/01/2026

Into the Omniverse: Physical AI Open Models and Frameworks Advance Robots and Autonomous Systems

Editor's note: This post is part of Into the Omniverse, a series focused on ...

29/01/2026

GeForce NOW Brings GeForce RTX Gaming to Linux PCs

Get ready to game - the native GeForce NOW app for Linux PCs is now available in beta, letting Linux desktops tap directly into GeForce RTX performance from the...

28/01/2026

Accelerating Science: A Blueprint for a Renewed National Quantum Initiative

Quantum technologies are rapidly emerging as foundational capabilities for economic competitiveness, national security and scientific leadership in the 21st cen...

22/01/2026

NVIDIA DRIVE AV Raises the Bar for Vehicle Safety as Mercedes-Benz CLA Earns Top Euro NCAP Award

AI-powered driver assistance technologies are becoming standard equipment, funda...

22/01/2026

Flight Controls Are Cleared for Takeoff on GeForce NOW

The wait is over, pilots. Flight control support - one of the most community-requested features for GeForce NOW - is live starting today, following its announce...

22/01/2026

From Pilot to Profit: Survey Reveals the Financial Services Industry Is Doubling Down on AI Investment and Open Source

AI has taken center stage in financial services, automating the research and exe...

22/01/2026

How to Get Started With Visual Generative AI on NVIDIA RTX PCs

AI-powered content generation is now embedded in everyday tools like Adobe and Canva, with a slew of agencies and studios incorporating the technology into thei...

21/01/2026

Largest Infrastructure Buildout In Human History': Jensen Huang on AI's Five-Layer Cake' at Davos

From skilled trades to startups, AI's rapid expansion is the beginning of th...

21/01/2026

Largest Infrastructure Buildout In Human History: Jensen Huang on AI's Five-Layer Cake at Davos

From skilled trades to startups, AI's rapid expansion is the beginning of th...

15/01/2026

Survive the Quarantine Zone and More With Devolver Digital Games on GeForce NOW

NVIDIA kicked off the year at CES, where the crowd buzzed about the latest gaming announcements - including the native GeForce NOW app for Linux and Amazon Fire...

13/01/2026

CEOs of NVIDIA and Lilly Share Blueprint for What Is Possible' in AI and Drug Discovery

NVIDIA and Lilly are putting together a blueprint for what is possible in the f...

09/01/2026

NVIDIA Unveils Multi-Agent Intelligent Warehouse and Catalog Enrichment AI Blueprints to Power the Retail Pipeline

Every that was easy shopping moment is made possible by teams working to hit s...

08/01/2026

Japan Science and Technology Agency Develops NVIDIA-Powered Moonshot Robot for Elderly Care

The next universal technology since the smartphone is on the horizon - and it ma...

08/01/2026

AI Copilot Keeps Berkeley's X-Ray Particle Accelerator on Track

In the rolling hills of Berkeley, California, an AI agent is supporting high-stakes physics experiments at the Advanced Light Source (ALS) particle accelerator....

08/01/2026

More Ways to Play, More Games to Love - GeForce NOW Wraps CES With Linux Support, Fire TV App, Flight Stick Controls

NVIDIA is wrapping up a big week at the CES trade show with a set of GeForce NOW...

07/01/2026

From Warehouse to Wallet: New State of AI in Retail and CPG Survey Uncovers How AI Is Rewiring Supply Chains and Customer Experiences

AI has transformed retail and consumer packaged goods (CPG) operations, enhancin...

05/01/2026

NVIDIA Expands Global DRIVE Hyperion Ecosystem to Accelerate the Road to Full Autonomy

At the CES trade show running this week in Las Vegas, NVIDIA announced that the ...

05/01/2026

NVIDIA DGX Spark and DGX Station Power the Latest Open-Source and Frontier Models From the Desktop

Open-source AI is accelerating innovation across industries, and NVIDIA DGX Spar...

05/01/2026

NVIDIA DGX SuperPOD Sets the Stage for Rubin-Based Systems

NVIDIA DGX SuperPOD is paving the way for large-scale system deployments built on the NVIDIA Rubin platform - the next leap forward in AI computing. At the CES...

05/01/2026

NVIDIA BlueField-Powered Cybersecurity and Acceleration Arrive on NVIDIA Enterprise AI Factory Validated Design

AI is powering breakthroughs across industries, helping enterprises operate with...

05/01/2026

NVIDIA Rubin Platform, Open Models, Autonomous Driving: NVIDIA Presents Blueprint for the Future at CES

NVIDIA founder and CEO Jensen Huang took the stage at the Fontainebleau Las Vega...

05/01/2026

NVIDIA DLSS 4.5, Path Tracing and G-SYNC Pulsar Supercharge Gameplay With Enhanced Performance and Visuals

At the CES trade show, NVIDIA today announced DLSS 4.5, which introduces Dynamic...

05/01/2026

NVIDIA RTX Accelerates 4K AI Video Generation on PC With LTX-2 and ComfyUI Upgrades

2025 marked a breakout year for AI development on PC. PC-class small language m...

05/01/2026

NVIDIA Brings GeForce RTX Gaming to More Devices With New GeForce NOW Apps for Linux PC and Amazon Fire TV

Announced at the CES trade show running this week in Las Vegas, NVIDIA is bringi...

01/01/2026

GeForce NOW Rings In 2026 With 14 New Games in January

New year, new games, all with RTX 5080-powered cloud energy. GeForce NOW is kicking off 2026 by looking back at an unforgettable year of wins and wildly high fr...

25/12/2025

Make Spirits Bright With Holiday Hits on GeForce NOW

Holiday lights are twinkling, hot cocoa's on the stove and gamers are settling in for a well-earned break. Whether staying in or heading on a winter getawa...

22/12/2025

Marine Biological Laboratory Explores Human Memory With AI and Virtual Reality

The works of Plato state that when humans have an experience, some level of change occurs in their brain, which is powered by memory - specifically long-term me...

18/12/2025

NVIDIA, US Government to Boost AI Infrastructure and R&D Investments Through Landmark Genesis Mission

NVIDIA will join the U.S. Department of Energy's (DOE) Genesis Mission as a ...

18/12/2025

Now Generally Available, NVIDIA RTX PRO 5000 72GB Blackwell GPU Expands Memory Options for Desktop Agentic AI

Top-notch options for AI at the desktops of developers, engineers and designers ...

18/12/2025

Deck the Vaults: Fallout: New Vegas' Joins the Cloud This Holiday Season

Step out of the vault and into the future of gaming with Fallout: New Vegas streaming on GeForce NOW, just in time to celebrate the newest season of the hit Ama...

17/12/2025

UC San Diego Lab Advances Generative AI Research With NVIDIA DGX B200 System

The Hao AI Lab research team at the University of California San Diego - at the forefront of pioneering AI model innovation - recently received an NVIDIA DGX B...