Run Google DeepMind's Gemma 3n on NVIDIA Jetson and RTX

26/06/2025

As of today, NVIDIA now supports the general availability of Gemma 3n on NVIDIA RTX and Jetson. Gemma, previewed by Google DeepMind at Google I/O last month, includes two new models optimized for multi-modal on-device deployment.

Gemma now includes audio in addition to the text and vision capabilities introduced in version 3.5. Each component integrates trusted research models: Universal Speech Model for audio, MobileNet v4 for vision, and MatFormer for text.

The biggest usage advancement is an innovation called Per-Lay Embeddings. It allows for significant reduction in RAM usage for parameters. The Gemma 3n E4B model has a raw parameter count of 8B parameters but can operate using a dynamic memory footprint that's comparable to a 4B model. This enables developers to use a higher quality model within a resource-constrained environment.

Model name Raw Parameters Input Context Length Output Context Length Size on Disk

E2B 5B 32K 32K subtracting request input 1.55GB

E4B 8B 32K 32K subtracting request input 2.82BB

Table 1: Gemma 3n model components for both the E2B and E4B model Powering robotics and edge AI with Jetson The Gemma family of models works well on NVIDIA Jetson devices that are geared at powering edge applications, such as next-generation robotics. The lightweight architecture and, now, dynamic memory usage fit in resource-constrained environments.

Jetson developers can participate in the Gemma 3n Impact Challenge hosted on Kaggle. The aim is to use this technology to create meaningful, positive change in the world in areas such as accessibility, education, healthcare, environmental sustainability, and crisis response. Several cash prizes, which start at $10,000, are available for submissions for overall placement and for using different technologies suited for on-device deployment, such as Jetson.

To get started, check out the live text and image demo from the Gemma 3 Developer Day in April and the GitHub repository for deploying Gemma locally using Ollama.

NVIDIA RTX for Windows developers and AI enthusiasts With NVIDIA RTX AI PCs, developers can easily deploy Gemma 3n models using Ollama. AI enthusiasts can use Gemma 3n models with RTX accelerations in their favorite apps like AnythingLLM and LM Studio.

Developers can deploy Gemma 3n locally to both RTX and Jetson devices with a few simple instructions using the Ollama CLI:

Download and install Ollama for Windows

Open a terminal window and complete the following commands:

ollama pull gemma3n:e4b ollama run gemma3n:e4b Summarize Shakespeare's Hamlet

NVIDIA collaborates with Ollama to provide performance optimizations for NVIDIA RTX GPUs, accelerating the latest models like Gemma 3n. For this model, Ollama leverages the Ollama engine in the backend, which builds upon the GGML library. Learn more about NVIDIA's contributions to the GGML library for maximum performance on NVIDIA RTX GPUs.

Customize Gemma for your data with the open NVIDIA NeMo Framework Developers can use the Gemma 3n models from Hugging Face with the open source NVIDIA NeMo Framework. It provides a comprehensive framework for post-training Llama models to achieve higher accuracy, specifically through fine-tuning with enterprise-specific data. The workflow within NeMo is designed to be end-to-end, covering data preparation, efficient fine-tuning, and model evaluation.

data-src=https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-png.webp alt=A diagram showing the workflow of NeMo Framework. It provides end-to-end support for developing large language models (LLMs) and multimodal models (MMs). class=lazyload wp-image-102645 data-srcset=https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-png.webp 1600w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-300x169-png.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-625x352-png.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-179x101-png.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-768x432-png.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-1536x864-png.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-645x363-png.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-660x370-png.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-500x281-png.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-160x90-png.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-362x204-png.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-196x110-png.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-1024x576-png.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2025/06/powerade-fig-1-960x540-png.webp 960w data-sizes=(max-width: 1600px) 100vw, 1600px />

Figure 1. NeMo Framework provides end-to-end support for large language models and multimodal models.

The workflow includes:

Data curation (NeMo Curator): Curator prepares high-quality datasets for either pretraining or fine-tuning by offering tools to extract, filter, and deduplicate large volumes of structured and unstructured data. It ensures the quality of the input data for the model.

Fine-tuning (NeMo): Once the data is curated, NeMo enables efficient fine-tuning of Llama models. It supports various techniques to optimize this process, including LoRA (Low-Rank Adaptation), PEFT (Parameter-Efficient Fine-Tuning), and full parameter tuning for comprehensive customization.

Model evaluation (NeMo Evaluator): After fine-tuning, NeMo Evaluator is used

LINK:	https://developer.nvidia.com/blog/run-google-deepminds-gemma-3n-on-nvi...
	See more stories from nvidia

Run Google DeepMind's Gemma 3n on NVIDIA Jetson and RTX

More from Nvidia

28/05/2026

28/05/2026

26/05/2026

21/05/2026

21/05/2026

19/05/2026

18/05/2026

14/05/2026

13/05/2026

13/05/2026

12/05/2026

07/05/2026

07/05/2026

06/05/2026

05/05/2026

30/04/2026

30/04/2026

28/04/2026

28/04/2026

23/04/2026

23/04/2026

22/04/2026

20/04/2026

20/04/2026

16/04/2026

15/04/2026

15/04/2026

09/04/2026

02/04/2026

02/04/2026

31/03/2026

26/03/2026

26/03/2026

25/03/2026

25/03/2026

24/03/2026

23/03/2026

19/03/2026

17/03/2026

17/03/2026

17/03/2026

12/03/2026

12/03/2026

11/03/2026

10/03/2026

10/03/2026

10/03/2026

10/03/2026

09/03/2026

09/03/2026