Sony Pixel Power calrec Sony

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

24/08/2026

The next era of AI inference won't be defined by a single breakthrough chip, network or system. It'll be defined by how every layer of the AI factory works together. That's why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform.

Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.

As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems.

https://blogs.nvidia.com/wp-content/uploads/2026/08/HotChips_DionHarris_v5_28MB.mp4

These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale.

At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.

Extreme Codesign Optimizes for Performance Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences.

NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.

NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.

Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA's full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.

https://blogs.nvidia.com/wp-content/uploads/2026/08/HotChips2026_Sizzle_final_30MB.mp4

Tuesday, Aug. 24, 8:00 a.m. PT

NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications.

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.

Tuesday, Aug. 24, 8:00 a.m. PT

SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation.

The SpaceXAI partnership extends NVIDIA's full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.

Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.

Tuesday, Aug. 24, 8:00 a.m. PT

NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.

Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation.

NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency.

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.

Building the Token Fact
LINK: https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/...
See more stories from nvidia

More from Nvidia

08/10/2026

Rally Up: Gears of War: E-Day' Launches on GeForce NOW

Gears of War: E-Day leads the charge on GeForce NOW this week, bringing Marcus Fenix and Dom Santiago's first fight against the Locust Horde to the cloud wi...

07/10/2026

NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs with RTX Spark and AI Agents

At a Microsoft event in San Francisco on Wednesday, Jensen Huang and Satya Nadel...

06/10/2026

Why Telecom Operators Are Building Their AI Strategy on Open Models

Telecom operators are increasingly building their AI strategies on open models - and the reasons go beyond mere cost. Open models give telcos the ability to t...

05/10/2026

From Scan to Treatment Plan, AI Helps Close Breast Cancer's Deadliest Gaps

Breast cancer is the most commonly diagnosed cancer among American women - yet the gaps in care are wide. A majority of women over age 40 skip the recommended a...

02/10/2026

NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI

Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to ...

01/10/2026

How NVIDIA GPUs Help Accelerate OpenAI's GPT-6 Astra Ultrafast

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by infer...

01/10/2026

Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment

AI factories are built by the megawatt, even by the gigawatt. Each megawatt fact...

01/10/2026

Fall Into 25 New Games on GeForce NOW This October

Spooky season is streaming in. Alongside falling leaves, pumpkin spice and everything nice, 25 new games are joining GeForce NOW throughout October, including s...

30/09/2026

NVIDIA Opens Applications for 2027-2028 Graduate Fellowships With Awards Up to $60,000

Bringing together the world's brightest minds and the latest accelerated com...

30/09/2026

From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI

Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that's still re...

24/09/2026

How Open Science Can Help Researchers Prepare for the Next Pandemic

When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus' key proteins well eno...

24/09/2026

Contain the Chaos: CONTROL Resonant' Launches on GeForce NOW

A warped Manhattan is waiting in the cloud this week. Remedy Entertainment's CONTROL Resonant brings Dylan Faden's extraordinary abilities and a paranat...

23/09/2026

Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale

When Sakeena Fiza describes her work as a validation engineer at NVIDIA, she does so in terms more befitting a detective story than a world-class engineering la...

22/09/2026

At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia

NVIDIA AI Day Singapore, which takes place Sept. 22-23 at the Raffles City Conve...

22/09/2026

NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development

To build and deploy sophisticated robotics applications that can perceive, reason and act in dynamic environments, developers need new physical AI models and to...

21/09/2026

Why Deploying Physical AI at Scale Demands Safety at Every Layer

Physical AI is moving rapidly from research to large-scale deployment. By 2035, ABI Research projects an installed base of 49 million level 3-5 autonomous vehic...

21/09/2026

From Enablement to Execution, Egypt's AI Ecosystem Reaches Production Scale

Today, Egypt's AI builders gathered in the Grand Egyptian Museum for a reception that highlighted the nation's rapidly growing AI ecosystem - spanning A...

21/09/2026

NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories

Every AI factory needs power and cooling that fit its computing architecture. As AI infrastructure expands, power, cooling, water, site and grid constraints are...

21/09/2026

AI Security Is an Engineering Problem - How to Solve It at Every Layer of the Agent Stack

AI security is an engineering problem. That means defined security requirements,...

21/09/2026

5 Companies Using NVIDIA AI for Clean Energy

Clean energy isn't hard to come by, but the pace of large-scale adoption has historically been slow due to bottlenecks - including out-of-date infrastructur...

17/09/2026

Cute Critters Come to the Cloud: Aniimo' Launches on GeForce NOW

A new creature-catching adventure is ready to stream from the cloud this week. Pawprint Studio's Aniimo arrives on GeForce NOW at launch, inviting gamers to...

16/09/2026

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

System performance, efficient infrastructure scaling and continuous software opt...

16/09/2026

Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers

AI factories are the infrastructure of the intelligence era. Scaling them respon...

15/09/2026

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to a...

15/09/2026

Now We Can Know Everything and Do Anything,' Jensen Huang Says at Dreamforce

Know everything. Do anything. That was the message NVIDIA founder and CEO Jensen Huang brought to Salesforce Dreamforce Tuesday, joining CEO Marc Benioff onsta...

15/09/2026

University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK

Air pollution is a serious public health risk, contributing to an estimated 30,0...

14/09/2026

Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX

As local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information on the device. Portable Computer is a ...

10/09/2026

Physical AI Takes the Wheel: How the World's Robotaxi Leaders Are Building With NVIDIA Technologies

The global robotaxi market - physical AI's first commercial breakthrough - i...

10/09/2026

Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

Manufacturing floors, warehouses and production lines rarely stay fixed - tasks change, layouts shift and new products arrive, and most robots can't keep up...

10/09/2026

Boots on the Ground: WARDOGS' Goes All Out on GeForce NOW at Early-Access Launch

Gear up: The latest PC games and major updates are ready to play on GeForce NOW ...

10/09/2026

d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

AI inference chipmaker d-Matrix today announced it will use NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform ...

09/09/2026

NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC

At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming together to turn ideas into action and dis...

03/09/2026

Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents ...

03/09/2026

NBA 2K27' With NVIDIA DLSS 5 Leads 26 New Games Coming to GeForce NOW

September is here with 28 more games streaming on GeForce NOW this month, led by a slam dunk: NBA 2K27 with the NVIDIA DLSS 5 3D-Guided Neural Rendering feature...

01/09/2026

NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier

We're at an inflection point in cybersecurity, Jensen Huang told a sold-out crowd at CrowdStrike's Fal.Con 2026 in Las Vegas Tuesday. Attacks are now a...

27/08/2026

GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

NVIDIA's Gamescom announcements are revealing what's next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more bi...

26/08/2026

NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastru...

25/08/2026

Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark

NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technolo...

24/08/2026

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

According to OpenRouter data, agentic AI workloads consume 15x more tokens than ...

24/08/2026

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

The next era of AI inference won't be defined by a single breakthrough chip,...

24/08/2026

How XPUs Meet a World-Class AI Factory

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost ...

20/08/2026

Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support

It's a new way into the cloud. GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight ...

14/08/2026

Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia's First University AI Center to Develop Local AI Talent

Indonesia is taking charge of its AI future. This week, the Ministry of Communi...

13/08/2026

Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More

GeForce NOW is giving cloud gaming an extra-credit upgrade just in time for back-to-school season. The native Linux app for GeForce NOW is officially out of be...

12/08/2026

NVIDIA CEO Tops Glassdoor's 2026 List of Best CEOs

NVIDIA founder and CEO Jensen Huang is ranked No. 1 on Glassdoor's Best CEOs list for 2026. In the just-released ranking, recognition is earned directly fr...

11/08/2026

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

We announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobiliz...

11/08/2026

Why Scaling AI Compute Performance Requires a New Power Architecture

Every new generation of accelerated computing demands more from the infrastructure underneath it - more compute performance, higher rack density and more effici...

11/08/2026

NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

As AI shifts from chatbots to autonomous agents, open models are serving market ...

11/08/2026

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout Au...