
Today, Google DeepMind released DiffusionGemma - an experimental open model built for exceptionally fast text generation. NVIDIA has optimized DiffusionGemma to run even faster across NVIDIA GeForce RTX GPUs, the NVIDIA RTX PRO platform and NVIDIA DGX Spark systems, from local PCs to the cloud.
Rather than generating text one word at a time, DiffusionGemma generates multiple words in parallel to output whole blocks of text, opening a new, low-latency frontier for the kind of single-user workloads that developers, researchers and AI enthusiasts run every day.
Features of the new model include:
Parallel generation: DiffusionGemma denoises up to 256 tokens per step instead of predicting one at a time.
Built on Gemma 4: DiffusionGemma is built on Gemma 4, a 26-billion-parameter mixture-of-experts model that activates just 3.8 billion parameters per step, pairing a diffusion head with Google's Gemma 4 architecture.
Up to 4x faster performance: The boost means fast text generation, where single-user generation usually stalls - on local hardware.
Open and local: DiffusionGemma is open weights under a permissive Apache 2.0 license and runs entirely on RTX and DGX Spark - no cloud, no per-token cost - with day-zero support in Hugging Face Transformers, vLLM and Unsloth.
A Different Way to Generate Text Almost every large language model (LLM) in wide use today is autoregressive - meaning it generates text one token at a time, with each new word depending on the one before it. That sequential process is what makes interactive AI feel like it's typing.
DiffusionGemma takes a different path. Built on the Gemma 4 26B mixture-of-experts architecture, it generates text the way diffusion models generate images: by starting from noise and refining a whole block of text at once. Each step denoises up to 256 tokens in parallel rather than emitting a single token and waiting to compute the next.
The result is a model that thinks in blocks instead of sequentially. For latency-sensitive, single-user work - such as interactive chat, agentic loops or on-device assistants that plan and act - that parallelism translates into responses fast enough to keep pace with how developers think and iterate.
DiffusionGemma Flies on NVIDIA GPUs Generating one token at a time is fundamentally a memory-bound problem - a traditional LLM spends most of its time waiting on memory bandwidth, not doing math, which leaves a lot of compute on the table.
Diffusion flips the equation. Pulling a full 256-token block through the transformer in parallel is a compute-bound workload - exactly what NVIDIA GPUs are built for. NVIDIA Tensor Cores accelerate the dense parallel math, and the CUDA software stack lets the model run efficiently from day one without bespoke tuning. In short, the model's design plays directly to the GPU' s strengths.
That shows up in the numbers. DiffusionGemma delivers 1,000 tokens/sec on a single NVIDIA H100 Tensor Core GPU, 150 tokens/sec on NVIDIA DGX Spark and up to 2,000 tokens/sec on NVIDIA DGX Station - roughly 4x faster than an equivalent autoregressive model running in the same single-user regime.
That advantage holds across NVIDIA's full lineup, running:
Locally on the NVIDIA DGX Spark deskside personal AI supercomputer - powered by the NVIDIA GB10 Grace Blackwell Superchip with 128GB of unified memory - with the preinstalled NVIDIA AI software stack ready for prototyping, fine-tuning and fully local agent workflows.
On NVIDIA RTX PRO 6000 workstations, providing developers, researchers and AI professionals with the headroom to run local low-latency generation and agentic loops as part of a professional workflow.
On DGX Station, delivering best-in-class, local high-speed inference with up to 2,000 tokens/sec for low-latency text generation and agentic loops with 748GB of coherent memory.
On GeForce RTX GPUs, with llama.cpp support coming soon.
Get Started Locally The fastest way to start testing and prototyping the model is through Hugging Face Transformers, which runs DiffusionGemma on a GeForce RTX 5090 or DGX Spark out of the box. For higher-throughput inference, vLLM provides day-zero serving support.
For adapting the model to a specific task or domain, fine-tuning is available through Unsloth and NVIDIA NeMo framework, with ready-made DGX Spark playbooks to get a local environment running quickly. Check out the vLLM playbooks for DGX Spark , RTX PRO and DGX Station.
Try Diffusion Gemma on Hugging Face or test it for free using NVIDIA-hosted application programming interfaces at build.nvidia.com.
Go deeper on the architecture and local deployment by reading the NVIDIA technical blog and the Google DeepMind announcement.
#ICYMI: The Latest From RTX AI Garage NVIDIA researchers released SANA-WM, an open source world model that turns a single image and a camera path into a minute-long, 720p video with precise 6-DoF control. At just 2.6 billion parameters, its distilled version generates a full 60-second clip in 34 seconds on a single NVIDIA GeForce RTX 5090 GPU using the NVFP4 format - delivering up to 36x higher throughput than comparable open models while running on one GPU. Read the paper.
Building Windows agents just got a full toolset - NVIDIA and Microsoft rolled out turnkey agent sandboxing on native Windows - Microsoft eXecution Containers plus the NVIDIA OpenShell runtime - alongside up to 2x faster agentic inference and native Windows support for Hermes Agent.
DGX Spark goes from unboxing to a running agent in minutes - A streamlined NVIDIA NemoClaw install gets developers to a working local agent fast, with Qwen3.6-35B running up to 2.6x faster on vLLM. And the new cluster assistant in NVIDIA Sync links up to four DGX Spark units into one 512GB pool - enough for 400-billion-parameter models.
Plug in to RTX Spark on Fac
North America Stories
17/07/2026
In-venue and creative video staffers at the professional and collegiate level ha...
17/07/2026
Production workers at Brooklyn Bowl's Williamsburg location voted 15-1 to join IATSE Local 4. The bargaining unit covers 24 production workers at the venue,...
17/07/2026
DAZN and ADI Predictstreet have announced an exclusive global strategic partners...
17/07/2026
Zixi and Comcast Technology Solutions (CTS) have announced a strategic integrati...
17/07/2026
Professional Fighters League (PFL) has announced a multi-year partnership with E...
17/07/2026
Spectrum Business has announced Spectrum TV Control Pro, a centralized app-based...
17/07/2026
Clark Wire and Cable has announced that Rick Fernandez, Managing Director of Axxion Consulting, will serve as Independent Manufacturers Representative for Centr...
17/07/2026
TikTok, the NBA, and the WNBA have announced a multi-year global content partnership covering highlights distribution, creator access to marquee events, live-ga...
17/07/2026
Company alleges article contained false and misleading claims regarding customer data...
17/07/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
17/07/2026
(L-R) Edward James Olmos, Luis Valdez, Lou Diamond Phillips and Lupe Valdez attend American Pachuco: The Legend Of Luis Valdez Premiere during the 2026 Sundan...
17/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/07/2026
Foundry Releases SmartRoto for Nuke
Brie Clayton July 17, 2026
0 Comments
Spline-based AI powered plugin accelerates time-consuming rotoscoping, helping...
17/07/2026
A Short Documentary About a Giant Pencil Edited with DaVinci Resolve Studio
Brie Clayton July 17, 2026
0 Comments
SBIFF jury award winner finished fro...
17/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/07/2026
Lightware has published its second voluntary Sustainability Report, showing how energy efficiency, product longevity and responsible material use are increasing...
17/07/2026
Creative software developer Foundry today announced Griptape Enterprise, a new tier of the Griptape AI workflow orchestration platform, extended to meet the str...
17/07/2026
Creative software developer Foundry today announced the availability of SmartRoto, a new AI-powered plugin upgrade for Nuke, NukeX, Nuke Studio, and Nuke Indie ...
17/07/2026
The Final Release of Blackmagic Fairlight Live is Now Available!
Brie Clayton July 16, 2026
0 Comments
Fairlight Live is a new audio mixer designed fo...
17/07/2026
Leveraging All of Entertainment's IP
Andy Marken July 16, 2026
0 Comments
Nobody ever wins the games. Period. There are survivors. There\s no win...
17/07/2026
Think of a professional athlete. What separates elite performers is what happens...
16/07/2026
This recent graduate and Georgia native has turned an early fascination with live production into a growing passion for camera work
In the live-sports-video in...
16/07/2026
TNDV is marking the first anniversary of Aspiration 35, a mobile production truck built around ARRI Alexa 35 Live camera systems. Over its first year, the truck...
16/07/2026
ATSC has completed a major revision to its A/85 Recommended Practice: Techniques...
16/07/2026
FloSports has announced an exclusive global media partnership with CrossFit for the 2026 CrossFit Games, presented by Air National Guard, beginning July 21 in S...
16/07/2026
Sennheiser has released firmware version 1.4 for its Spectera wireless system and announced the Spectera Command Button (Cat. No. 701014). The update adds comma...
16/07/2026
ESPN and the Southwestern Athletic Conference (SWAC) have announced a multi-year...
16/07/2026
Audio-Technica has announced the 20 Series Control Application, a Stream Deck plug-in for compatible Audio-Technica USB microphones. The application is compatib...
16/07/2026
The fifteenth season of The Voice - La plus belle voix on TF1 used four Solid St...
16/07/2026
Gray Media and the Atlanta Hawks have announced a broadcast partnership that wil...
16/07/2026
Season 3 of Netflix's docuseries, which dropped this week, follows Jayden Da...
16/07/2026
The Texas Rangers and Rangers Sports Network presented by Progressive (RSN) have announced BZZR as the new direct-to-consumer distributor for Rangers game broad...
16/07/2026
With England's FIFA World Cup 2026 team getting ready to fly home after losi...
16/07/2026
In a return to Royal Birkdale, the team enjoys partnering with ETP and Gravity M...
16/07/2026
Along with a new studio at Ball Arena, the Colorado-based RSN is working with XR...
16/07/2026
Austria was back at the World Cup for the first time in 28 years, and the team d...
16/07/2026
With 42 RF cameras and high-power radio mics deployed, the spectrum environment ...
16/07/2026
June brought stabilization to the television market. Poles spent an average of 3 hours and 36 minutes a day in front of their TV screens exactly the same as i...
16/07/2026
Glensound returns to IBC with a major new addition to its intelligent loudspeaker range, the Greater Divine Supernal studio monitor, alongside recent developmen...
16/07/2026
33 leading exhibitors are taking part in the GREAT Britain and Northern Ireland Pavilions and linked locations across IBC2026 (Amsterdam RAI, 11 14 September)...
16/07/2026
Hitomi Broadcast will introduce Spectra, a new HDR colour verification solution measured in picture, not using VPID, at IBC2026 (Hall 10, Stand A40). Expanding ...
16/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
16/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
16/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
16/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
16/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
16/07/2026
Zixi, a leader in live video delivery and workflow orchestration, and Comcast Technology Solutions (CTS) today announced a strategic new integration to deliver ...
16/07/2026
Blackmagic Design Cameras Capture Immersive Flight for New Cirrus App
Brie Clayton July 15, 2026
0 Comments
Leader in Personal Aviation taps URSA Cin...
16/07/2026
XenData Announces LTO Archive Appliances with both File and S3 Object Storage In...