
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech.
Before a packed audience - with more than 8,000 attendees this year, up from 3,500 last year - Buck discussed new collaborations across NVIDIA platforms and more.
Amazon's Annapurna Labs is working with NVIDIA on the NVHBM custom high-bandwidth memory technology.
d-Matrix is integrating with NVLink Fusion to combine NVIDIA Vera CPUs with d-Matrix Raptor XPUs to deliver ultra low-latency inference at scale.
NVIDIA and partners unveiled new results as well:
Emerald AI and NVIDIA demonstrated a commercial AI factory flexible-load program, working with Silicon Valley Power.
Lambda improved performance per watt by 23% with NVIDIA DSX MaxLPS.
Pinterest is using the NVIDIA Blackwell platform and NVIDIA Dynamo inference software to bring conversational AI to visual discovery.
The news comes as agentic AI is driving a new class of workloads that demand more performance, efficiency and scale from AI infrastructure.
NVIDIA addresses that challenge with a full-stack AI factory platform spanning Vera Rubin systems, Dynamo inference software, NeMo libraries and NVIDIA networking - including NVIDIA NVLink for scale-up computing, Spectrum-X Ethernet and ConnectX SuperNICs for connecting thousands of nodes, BlueField-powered context-memory storage and BlueField DPUs for infrastructure security.
The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt. AI factories must now be codesigned from silicon to grid. NVIDIA DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink helps unite large-scale accelerated computing into a single high-performance system.
The result is AI infrastructure designed to generate more tokens, improve efficiency and help customers get more value from every megawatt of power.
Tuesday, Sept. 15, 9:55 a.m. PT
AI Factory Flexible-Load Program Unlocks More Grid Power Silicon Valley Power operates a flexible-load interconnection program that enables AI factories to support grid flexibility.
Through this program, Emerald AI worked with NVIDIA to demonstrate automated load reduction at Silicon Valley Power. The system successfully responded to hundreds of demand signals from Silicon Valley Power while protecting AI workload performance.
NVIDIA partner Emerald AI is planning to use NVIDIA DSX Flex for its Conductor grid-responsive power management software that dynamically adjusts an AI factory's energy consumption based on real-time electricity grid signals and hybrid energy sources.
Using DSX Flex to autonomously load balance, Emerald AI can demonstrate its ability to automatically throttle back power on low-priority AI jobs and then return to normal.
DSX Flex can receive grid signals - load-shedding requests, demand-response events and pricing signals - and automatically act within a predefined workload hierarchy. The most critical jobs keep going, while everything else pauses temporarily and then resumes. In this way, facilities can participate in energy efficiency for the grid.
This capability enables AI factories to operate as flexible grid resources, so they can reduce demand when the grid needs relief, protect priority AI workloads and prove that controllable AI load can help unlock more grid capacity for growth.
Learn more about the Emerald AI partnership with Silicon Valley Power.
Tuesday, Sept. 15, 9:55 a.m. PT
Lambda Maximizes Performance per Watt With NVIDIA DSX MaxLPS AI cloud provider Lambda released results at the AI Infra Summit, providing the first validation of NVIDIA DSX MaxLPS on NVIDIA Blackwell servers.
DSX MaxLPS continuously monitors power consumption across GPUs and racks, dynamically shifting available power where it's needed most and reclaiming capacity that static provisioning leaves unused. Because training and inference workloads have different power profiles, MaxLPS optimizes power allocation across mixed-workload AI factories.
The results were significant: Lambda ran 19 nodes within the same power budget typically allocated to 16 full-power nodes, increasing cluster-wide token throughput by 24% - from about 4 million to 5 million tokens per second - while improving performance per watt by 23%. For next-generation NVIDIA Vera Rubin NVL72 AI factories, DSX MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget in the right deployment environments.
This means AI factory operators can increase AI capacity and token throughput within the same power envelope, helping maximize the productivity and economic value of every megawatt.
Learn more about Lambda's results.
Tuesday, Sept. 15, 9:55 a.m. PT
Vera Rubin and Groq 3 LPX Turn More Power Into Tokens In AI factories, power is the constraint. Every watt counts, so squeezing more tokens from every megawatt is what matters most for AI infrastructure.
NVIDIA does this with a full-stack AI factory built on the Vera Rubin NVL72 - systems, networking, software and power management working as one. At the factory level, NVIDIA DSX MaxLPS dynamically shifts power across racks as demand rises and falls.
The payoff is significant:
Up to 40% more GPUs within the same site-power envelope
Up to 35% higher token throughput - no new power lines required
Inside each rack, Intelligent Power Smoothing software and expanded energy buffering absorb short spikes, letting systems run closer to sustained demand. Stranded headroom becomes productive compute.
For agentic AI - where agents chain reasoning steps and tool calls - latency and context length compounds quickly. That's where NVIDIA Gro
More from Nvidia
15/09/2026
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA,...
15/09/2026
On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to a...
15/09/2026
Know everything. Do anything.
That was the message NVIDIA founder and CEO Jensen Huang brought to Salesforce Dreamforce Tuesday, joining CEO Marc Benioff onsta...
15/09/2026
Air pollution is a serious public health risk, contributing to an estimated 30,0...
14/09/2026
As local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information on the device.
Portable Computer is a ...
10/09/2026
The global robotaxi market - physical AI's first commercial breakthrough - i...
10/09/2026
Manufacturing floors, warehouses and production lines rarely stay fixed - tasks change, layouts shift and new products arrive, and most robots can't keep up...
10/09/2026
Gear up: The latest PC games and major updates are ready to play on GeForce NOW ...
10/09/2026
AI inference chipmaker d-Matrix today announced it will use NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform ...
09/09/2026
At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming together to turn ideas into action and dis...
03/09/2026
Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents ...
03/09/2026
September is here with 28 more games streaming on GeForce NOW this month, led by a slam dunk: NBA 2K27 with the NVIDIA DLSS 5 3D-Guided Neural Rendering feature...
01/09/2026
We're at an inflection point in cybersecurity, Jensen Huang told a sold-out crowd at CrowdStrike's Fal.Con 2026 in Las Vegas Tuesday. Attacks are now a...
27/08/2026
NVIDIA's Gamescom announcements are revealing what's next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more bi...
26/08/2026
The next wave of AI is placing new demands on infrastructure.
As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastru...
25/08/2026
NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technolo...
24/08/2026
According to OpenRouter data, agentic AI workloads consume 15x more tokens than ...
24/08/2026
The next era of AI inference won't be defined by a single breakthrough chip,...
24/08/2026
To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost ...
20/08/2026
It's a new way into the cloud.
GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight ...
14/08/2026
Indonesia is taking charge of its AI future.
This week, the Ministry of Communi...
13/08/2026
GeForce NOW is giving cloud gaming an extra-credit upgrade just in time for back-to-school season.
The native Linux app for GeForce NOW is officially out of be...
12/08/2026
NVIDIA founder and CEO Jensen Huang is ranked No. 1 on Glassdoor's Best CEOs list for 2026.
In the just-released ranking, recognition is earned directly fr...
11/08/2026
We announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobiliz...
11/08/2026
Every new generation of accelerated computing demands more from the infrastructure underneath it - more compute performance, higher rack density and more effici...
11/08/2026
As AI shifts from chatbots to autonomous agents, open models are serving market ...
11/08/2026
The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally.
Throughout Au...
08/08/2026
The global buildout of AI infrastructure reached a new milestone today - Firebird, an emerging AI cloud, launched the CIS region's largest AI factory in Arm...
06/08/2026
Editor's note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners and enterprises can transform their workflows u...
06/08/2026
August is here, bringing 26 new games for GeForce NOW members.
Command the seas in World of Warships: Legends and discover what's next in the GeForce NOW ...
04/08/2026
Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory.
But rising needs aren't me...
04/08/2026
For robotaxis and other autonomous vehicles (AVs), the hardest problems aren'...
04/08/2026
Members of the Open Secure AI Alliance - now more than 120 organizations strong - are developing new guidelines to strengthen agentic AI cybersecurity as the an...
30/07/2026
Back to school means balancing assignments, deadlines and downtime. GeForce NOW makes it easy to have it all.
With cloud gaming, everyday laptops used for clas...
28/07/2026
As a discerning AI investor who values style and substance, Sarah Guo knows this...
27/07/2026
Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government...
26/07/2026
The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems. To help meet t...
23/07/2026
At this week's AI Summit in San Francisco, South Korean President Jae Myung Lee and some of the country's top business leaders and researchers are meeti...
23/07/2026
Lock in and load up the cloud. GFN Thursday brings fresh updates and new adventu...
22/07/2026
NVIDIA founder and CEO Jensen Huang today visited the Naval Postgraduate School in Monterey, California, to commission an NVIDIA DGX GB300 system - bringing one...
22/07/2026
Before a healthcare robot can be useful in the real world, it has to learn how the physical world pushes back. Anatomy varies. Instruments bend, press, slip and...
21/07/2026
The AI era runs on AI infrastructure. Many of these advanced systems are built a...
21/07/2026
AI has entered the gigascale era.
The world's most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier mod...
21/07/2026
NVIDIA Vera Rubin is here, and it's going gigascale.
Vera Rubin NVL72 produ...
20/07/2026
At this year's SIGGRAPH conference, running through Thursday, July 23, in Lo...
20/07/2026
Erin Davis calls it the SuperDuperPOD. That's two things in one name: phar...
17/07/2026
Think of a professional athlete. What separates elite performers is what happens...
16/07/2026
Onimusha: Way of the Sword is coming to GeForce NOW at launch, with the playable...
15/07/2026
General-purpose robots and autonomous machines are moving from research labs to ...
15/07/2026
Home to leading manufacturers, robotics pioneers, infrastructure builders and iconic gaming companies, of course, Japan is one of the world's centers of AI ...