Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech.Before a packed audience - with more than 8,000 attendees this year, up from 3,500 last year - Buck discussed new collaborations across NVIDIA platforms and more.
Amazon's Annapurna Labs is working with NVIDIA on the NVHBM custom high-bandwidth memory technology.
d-Matrix is integrating with NVLink Fusion to combine NVIDIA Vera CPUs with d-Matrix Raptor XPUs to deliver ultra low-latency inference at scale.
NVIDIA and partners unveiled new results as well:
Emerald AI and NVIDIA demonstrated a commercial AI factory flexible-load program, working with Silicon Valley Power.
Lambda improved performance per watt by 23% with NVIDIA DSX MaxLPS.
Pinterest is using the NVIDIA Blackwell platform and NVIDIA Dynamo inference software to bring conversational AI to visual discovery.
The news comes as agentic AI is driving a new class of workloads that demand more performance, efficiency and scale from AI infrastructure.
NVIDIA addresses that challenge with a full-stack AI factory platform spanning Vera Rubin systems, Dynamo inference software, NeMo libraries and NVIDIA networking - including NVIDIA NVLink for scale-up computing, Spectrum-X Ethernet and ConnectX SuperNICs for connecting thousands of nodes, BlueField-powered context-memory storage and BlueField DPUs for infrastructure security.
The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt. AI factories must now be codesigned from silicon to grid. NVIDIA DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink helps unite large-scale accelerated computing into a single high-performance system.
The result is AI infrastructure designed to generate more tokens, improve efficiency and help customers get more value from every megawatt of power.
Tuesday, Sept. 15, 9:55 a.m. PT
AI Factory Flexible-Load Program Unlocks More Grid Power Silicon Valley Power operates a flexible-load interconnection program that enables AI factories to support grid flexibility.
Through this program, Emerald AI worked with NVIDIA to demonstrate automated load reduction at Silicon Valley Power. The system successfully responded to hundreds of demand signals from Silicon Valley Power while protecting AI workload performance.
NVIDIA partner Emerald AI is planning to use NVIDIA DSX Flex for its Conductor grid-responsive power management software that dynamically adjusts an AI factory's energy consumption based on real-time electricity grid signals and hybrid energy sources.
Using DSX Flex to autonomously load balance, Emerald AI can demonstrate its ability to automatically throttle back power on low-priority AI jobs and then return to normal.
DSX Flex can receive grid signals - load-shedding requests, demand-response events and pricing signals - and automatically act within a predefined workload hierarchy. The most critical jobs keep going, while everything else pauses temporarily and then resumes. In this way, facilities can participate in energy efficiency for the grid.
This capability enables AI factories to operate as flexible grid resources, so they can reduce demand when the grid needs relief, protect priority AI workloads and prove that controllable AI load can help unlock more grid capacity for growth.
Learn more about the Emerald AI partnership with Silicon Valley Power.
Tuesday, Sept. 15, 9:55 a.m. PT
Lambda Maximizes Performance per Watt With NVIDIA DSX MaxLPS AI cloud provider Lambda released results at the AI Infra Summit, providing the first validation of NVIDIA DSX MaxLPS on NVIDIA Blackwell servers.
DSX MaxLPS continuously monitors power consumption across GPUs and racks, dynamically shifting available power where it's needed most and reclaiming capacity that static provisioning leaves unused. Because training and inference workloads have different power profiles, MaxLPS optimizes power allocation across mixed-workload AI factories.
The results were significant: Lambda ran 19 nodes within the same power budget typically allocated to 16 full-power nodes, increasing cluster-wide token throughput by 24% - from about 4 million to 5 million tokens per second - while improving performance per watt by 23%. For next-generation NVIDIA Vera Rubin NVL72 AI factories, DSX MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget in the right deployment environments.
This means AI factory operators can increase AI capacity and token throughput within the same power envelope, helping maximize the productivity and economic value of every megawatt.
Learn more about Lambda's results.
Tuesday, Sept. 15, 9:55 a.m. PT
Vera Rubin and Groq 3 LPX Turn More Power Into Tokens In AI factories, power is the constraint. Every watt counts, so squeezing more tokens from every megawatt is what matters most for AI infrastructure.
NVIDIA does this with a full-stack AI factory built on the Vera Rubin NVL72 - systems, networking, software and power management working as one. At the factory level, NVIDIA DSX MaxLPS dynamically shifts power across racks as demand rises and falls.
The payoff is significant:
Up to 40% more GPUs within the same site-power envelope
Up to 35% higher token throughput - no new power lines required
Inside each rack, Intelligent Power Smoothing software and expanded energy buffering absorb short spikes, letting systems run closer to sustained demand. Stranded headroom becomes productive compute.
For agentic AI - where agents chain reasoning steps and tool calls - latency and context length compounds quickly. That's where NVIDIA Gro










