Sony Pixel Power calrec Sony

NVIDIA Accelerates Google DeepMind's DiffusionGemma for Local AI

10/06/2026

Today, Google DeepMind released DiffusionGemma - an experimental open model built for exceptionally fast text generation. NVIDIA has optimized DiffusionGemma to run even faster across NVIDIA GeForce RTX GPUs, the NVIDIA RTX PRO platform and NVIDIA DGX Spark systems, from local PCs to the cloud.

Rather than generating text one word at a time, DiffusionGemma generates multiple words in parallel to output whole blocks of text, opening a new, low-latency frontier for the kind of single-user workloads that developers, researchers and AI enthusiasts run every day.

Features of the new model include:

Parallel generation: DiffusionGemma denoises up to 256 tokens per step instead of predicting one at a time.

Built on Gemma 4: DiffusionGemma is built on Gemma 4, a 26-billion-parameter mixture-of-experts model that activates just 3.8 billion parameters per step, pairing a diffusion head with Google's Gemma 4 architecture.

Up to 4x faster performance: The boost means fast text generation, where single-user generation usually stalls - on local hardware.

Open and local: DiffusionGemma is open weights under a permissive Apache 2.0 license and runs entirely on RTX and DGX Spark - no cloud, no per-token cost - with day-zero support in Hugging Face Transformers, vLLM and Unsloth.

A Different Way to Generate Text Almost every large language model (LLM) in wide use today is autoregressive - meaning it generates text one token at a time, with each new word depending on the one before it. That sequential process is what makes interactive AI feel like it's typing.

DiffusionGemma takes a different path. Built on the Gemma 4 26B mixture-of-experts architecture, it generates text the way diffusion models generate images: by starting from noise and refining a whole block of text at once. Each step denoises up to 256 tokens in parallel rather than emitting a single token and waiting to compute the next.

The result is a model that thinks in blocks instead of sequentially. For latency-sensitive, single-user work - such as interactive chat, agentic loops or on-device assistants that plan and act - that parallelism translates into responses fast enough to keep pace with how developers think and iterate.

DiffusionGemma Flies on NVIDIA GPUs Generating one token at a time is fundamentally a memory-bound problem - a traditional LLM spends most of its time waiting on memory bandwidth, not doing math, which leaves a lot of compute on the table.

Diffusion flips the equation. Pulling a full 256-token block through the transformer in parallel is a compute-bound workload - exactly what NVIDIA GPUs are built for. NVIDIA Tensor Cores accelerate the dense parallel math, and the CUDA software stack lets the model run efficiently from day one without bespoke tuning. In short, the model's design plays directly to the GPU' s strengths.

That shows up in the numbers. DiffusionGemma delivers 1,000 tokens/sec on a single NVIDIA H100 Tensor Core GPU, 150 tokens/sec on NVIDIA DGX Spark and up to 2,000 tokens/sec on NVIDIA DGX Station - roughly 4x faster than an equivalent autoregressive model running in the same single-user regime.

That advantage holds across NVIDIA's full lineup, running:

Locally on the NVIDIA DGX Spark deskside personal AI supercomputer - powered by the NVIDIA GB10 Grace Blackwell Superchip with 128GB of unified memory - with the preinstalled NVIDIA AI software stack ready for prototyping, fine-tuning and fully local agent workflows.

On NVIDIA RTX PRO 6000 workstations, providing developers, researchers and AI professionals with the headroom to run local low-latency generation and agentic loops as part of a professional workflow.

On DGX Station, delivering best-in-class, local high-speed inference with up to 2,000 tokens/sec for low-latency text generation and agentic loops with 748GB of coherent memory.

On GeForce RTX GPUs, with llama.cpp support coming soon.

Get Started Locally The fastest way to start testing and prototyping the model is through Hugging Face Transformers, which runs DiffusionGemma on a GeForce RTX 5090 or DGX Spark out of the box. For higher-throughput inference, vLLM provides day-zero serving support.

For adapting the model to a specific task or domain, fine-tuning is available through Unsloth and NVIDIA NeMo framework, with ready-made DGX Spark playbooks to get a local environment running quickly. Check out the vLLM playbooks for DGX Spark , RTX PRO and DGX Station.

Try Diffusion Gemma on Hugging Face or test it for free using NVIDIA-hosted application programming interfaces at build.nvidia.com.

Go deeper on the architecture and local deployment by reading the NVIDIA technical blog and the Google DeepMind announcement.

#ICYMI: The Latest From RTX AI Garage NVIDIA researchers released SANA-WM, an open source world model that turns a single image and a camera path into a minute-long, 720p video with precise 6-DoF control. At just 2.6 billion parameters, its distilled version generates a full 60-second clip in 34 seconds on a single NVIDIA GeForce RTX 5090 GPU using the NVFP4 format - delivering up to 36x higher throughput than comparable open models while running on one GPU. Read the paper.

Building Windows agents just got a full toolset - NVIDIA and Microsoft rolled out turnkey agent sandboxing on native Windows - Microsoft eXecution Containers plus the NVIDIA OpenShell runtime - alongside up to 2x faster agentic inference and native Windows support for Hermes Agent.

DGX Spark goes from unboxing to a running agent in minutes - A streamlined NVIDIA NemoClaw install gets developers to a working local agent fast, with Qwen3.6-35B running up to 2.6x faster on vLLM. And the new cluster assistant in NVIDIA Sync links up to four DGX Spark units into one 512GB pool - enough for 400-billion-parameter models.

Plug in to RTX Spark on Fac
LINK: https://blogs.nvidia.com/blog/rtx-ai-garage-local-gemma-diffusion/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

07/10/2026

Dalet Flex LTS Delivers Smarter Media Operations from Ingest to Distribution

Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

21/07/2026

Savannah Bananas Tap Calrec to Expand Banana Ball TV Coverage

Share Copy link Facebook X Linkedin Bluesky Email...

21/07/2026

Where Growth Takes Root Viaccess-Orca Showcases the Found...

Across the TV and streaming ecosystem, industry players face mounting pressure to do more with less. They must simplify operations, protect premium content and ...

21/07/2026

Tuxera breaks performance ceilings and brings SMB/NFS mul...

Tuxera, a leading prover of quality-assured file systems and networking technologies, is bringing its latest advances in connectivity performance to the media w...

21/07/2026

COW Jobs: Apple Motion Expert for Custom FCP Rigged Graphics Package

COW Jobs: Apple Motion Expert for Custom FCP Rigged Graphics Package Brie Clayton July 21, 2026 0 Comments HIRING: Apple Motion Expert for Custom FCP ...

21/07/2026

Cinegy directly addresses the real issues of software-def...

Cinegy GmbH, the premier provider of software-defined television technology, will use its presence at IBC2026 (stand 7.A01, Amsterdam RAI, 11 14 September) to...

21/07/2026

Calrec brings leading audio solutions and long term busin...

Calrec will be located in Hall 8, on Stand C47 Beyond bigger For years, broadcast facilities were built around one assumption: provision for the biggest pro...

21/07/2026

Pebble takes the next step towards the future of media de...

Pebble, the leading automation, content management and integrated channel specialist, will discuss its future-facing developments for the new generation of medi...

21/07/2026

nxtedition Brings Production, AI and Automation Together...

nxtedition returns to IBC2026 with its consolidated production platform, showing how scripting, editing, graphics, AI-assisted tools and automation can work wit...

21/07/2026

Big Blue Marble brings C2PA Content Credentials to Cloud...

C2PA content signing helps broadcasters, public institutions, and publishers give audiences a verifiable record of where their video content came from and how i...

21/07/2026

HDHomeRun Enables Operation During Internet Outages

Share Copy link Facebook X Linkedin Bluesky Email...

21/07/2026

Calif. Federal Judge Pauses Paramount-WBD Merger

Share Copy link Facebook X Linkedin Bluesky Email...

21/07/2026

AWARN Rebuts Weigel Claims of 3.0 EAS Problems

Share Copy link Facebook X Linkedin Bluesky Email...

21/07/2026

FCC Announces Tentative Agenda for August Open Meeting

Share Copy link Facebook X Linkedin Bluesky Email...

21/07/2026

VIDA Introduces QC Manager for Collaborative Quality Cont...

New capability centralizes collaboration, feedback, security and status tracking within VIDA, providing a structured, auditable approach to master quality contr...

21/07/2026

Lightcraft Launches Exclusive Spark Story Beta for Filmmakers and Creators at SIGGRAPH 2026

Lightcraft Launches Exclusive Spark Story Beta for Filmmakers and Creators at SI...

21/07/2026

Indie Road Movie Where in the Hell Shot with Pocket Cinema Camera 4K

Indie Road Movie Where in the Hell Shot with Pocket Cinema Camera 4K Brie Clayton July 20, 2026 0 Comments Colorist blends vintage film looks to shape...

21/07/2026

Which USS Defiant Pulse Phaser effect is better?

Which USS Defiant Pulse Phaser effect is better? Graham Quince July 20, 2026 0 Comments Aargh, ever since @DarkRavenProductions posted a comment ask...

21/07/2026

Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories

AI has entered the gigascale era. The world's most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier mod...

21/07/2026

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

NVIDIA Vera Rubin is here, and it's going gigascale. Vera Rubin NVL72 produ...

21/07/2026

RT announces biggest ever year for new drama

160 hours of Irish storytelling for 2026 KIN and The Walsh Sisters return RT has today unveiled its most ambitious drama slate ever, with a record-breakin...

20/07/2026

SVG All-Stars: Rob Coons, Senior Director, StudentU, Big Ten Network

The former Northwestern broadcast-operations leader is helping train the next generation of live-sports-production talent across the Big Ten The sports-product...

20/07/2026

Give Me the Backstory: Get to Know Stacey Lee, the Filmmaker Behind Murder 101

By Lucy Spicer One of the most exciting things about the Sundance Film Festival is having a front-row seat for the bright future of independent filmmaking. Whi...

20/07/2026

Graph Tech Guitar Labs open UK Online Store

Product range now readily available in the UK Graph Tech Guitar Labs have just announced the launch of a new UK Online Store that makes it quicker and easie...

20/07/2026

Audeze launch the Maxwell 2 ANC

Popular headset gains active noise cancellation The Maxwell headet was Audeze's first foray into the gaming world, and thanks to its Dolby Atmos compati...

20/07/2026

The National Film and Video Foundation (NFVF) Call for Public Screening funding applications for Cycle 1, 2026/27 financial year is Open

The NFVF, an agency of the Department of Sport, Arts and Culture, has released t...

20/07/2026

Telemundo Inks Another Major U.S. Spanish-Language Soccer Deal

Share Copy link Facebook X Linkedin Bluesky Email...

20/07/2026

HDHomeRun Enables Operation During An Internet Outage

Share Copy link Facebook X Linkedin Bluesky Email...

20/07/2026

Starfish highlights flexible, scalable transport stream p...

Starfish Technologies will use IBC2026 to showcase the flexibility of its transport stream processing software, including the latest versions of TS Splicer (Win...

20/07/2026

Bitfocus makes the connections at IBC2026

Bitfocus, the specialist in media control and monitoring, will show at IBC2026 (Elgato stand 8.D31, Amsterdam RAI, 11 14 September) how its Buttons control la...

20/07/2026

Mediagenix Introduces Trusted Agentic AI Operating Model...

Mediagenix, a global leader in smart content solutions to profitably connect the right content to the right audience, today announced new AI capabilities that e...

20/07/2026

Big Blue Marble Cloud DRM nominated for Streaming Media R...

Big Blue Marble's Cloud DRM has been nominated in the DRM/Content Protection category of the 2026 Streaming Media Readers' Choice Awards. Only four pro...

20/07/2026

Sky brings free global roaming to millions of loyal customers when they choose Sky Mobile

Monday 20 July 2026 Sky brings free global roaming to millions of loyal custome...

20/07/2026

Red Seat Ventures Announces Introduction of Premium Creator and Podcast Communities to Amazon DSP

Red Seat Ventures Announces Introduction of Premium Creator and Podcast Communit...

20/07/2026

RT secures exclusive free-to-air Irish rights to the 2030 FIFA World Cup

FIFA World Cup Final sets new RT Player record as the most-streamed single event in the platform's history Over 1 million viewers watched live on RT 2 as ...

20/07/2026

At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI

At this year's SIGGRAPH conference, running through Thursday, July 23, in Lo...

20/07/2026

Bristol Myers Squibb Building Life Science Industry's Most Advanced AI Factory on NVIDIA Vera Rubin

Erin Davis calls it the SuperDuperPOD. That's two things in one name: phar...

19/07/2026

Halftime Show at FIFA World Cup Final Joins a Litany of Firsts for the Quadrennial Event

Justin Bieber, Madonna, Shakira, BTS make for a diverse lineup, and the venue ad...

19/07/2026

More Than Just a Game: FIFA World Cup's Lance Brass Breaks Down Stadium Production and Entertainment

The stage is set: three-time champion Argentina will defend its World Cup title ...

19/07/2026

Acustica reveal Mystic 2

Channel strip plug-in gets upgraded Acustica Audio's vintage-inspired channel strip plug-in has just been treated to an update that expands its tonal ra...

18/07/2026

More Than Just a Game: FIFA World Cups Lance Brass Breaks Down Stadium Production & Entertainment

Topics include pre-match ceremonies, live performances, the tournament's fir...

18/07/2026

As the Final Approaches, FIFA and HBS Take Stock of a World Cup That Rewrote the Production Playbook

When FIFA and HBS set out to produce the 2026 FIFA World Cup, the numbers alone ...

18/07/2026

IK Multimedia add Brown Panel Signature Collection to TONEX

Captures nine sought-after Fender amps IK Multimedia's latest TONEX expansion captures a selection of nine rare Brown Panel' Fender amps that were ...

18/07/2026

Frap Tools update the Magnolia

Latest batch ships alongside firmware update Since being unveiled at Superbooth 2025, Frap Tools' debut polysynth has been met with widespread praise, a...

18/07/2026

Netflix Viewing Hit Record 97 Billion Hours in First Half of 2026

Share Copy link Facebook X Linkedin Bluesky Email...

18/07/2026

YouTube's Creative Ecosystem Contributed $60 Billion to U.S. GDP

Share Copy link Facebook X Linkedin Bluesky Email...

17/07/2026

SVG GameDay, Ep. 24: Mercedes-Benz Stadiums Cole Gallagher - Supporting Shows in the ATL

In-venue and creative video staffers at the professional and collegiate level ha...