Lightweight Champ: NVIDIA Releases Small Language Model With State-of-the-Art Accuracy

21/08/2024

Developers of generative AI typically face a tradeoff between model size and accuracy. But a new language model released by NVIDIA delivers the best of both, providing state-of-the-art accuracy in a compact form factor.

Mistral-NeMo-Minitron 8B - a miniaturized version of the open Mistral NeMo 12B model released by Mistral AI and NVIDIA last month - is small enough to run on an NVIDIA RTX-powered workstation while still excelling across multiple benchmarks for AI-powered chatbots, virtual assistants, content generators and educational tools. Minitron models are distilled by NVIDIA using NVIDIA NeMo, an end-to-end platform for developing custom generative AI.

We combined two different AI optimization methods - pruning to shrink Mistral NeMo's 12 billion parameters into 8 billion, and distillation to improve accuracy, said Bryan Catanzaro, vice president of applied deep learning research at NVIDIA. By doing so, Mistral-NeMo-Minitron 8B delivers comparable accuracy to the original model at lower computational cost.

Unlike their larger counterparts, small language models can run in real time on workstations and laptops. This makes it easier for organizations with limited resources to deploy generative AI capabilities across their infrastructure while optimizing for cost, operational efficiency and energy use. Running language models locally on edge devices also delivers security benefits, since data doesn't need to be passed to a server from an edge device.

Developers can get started with Mistral-NeMo-Minitron 8B packaged as an NVIDIA NIM microservice with a standard application programming interface (API) - or they can download the model from Hugging Face. A downloadable NVIDIA NIM, which can be deployed on any GPU-accelerated system in minutes, will be available soon.

State-of-the-Art for 8 Billion Parameters For a model of its size, Mistral-NeMo-Minitron 8B leads on nine popular benchmarks for language models. These benchmarks cover a variety of tasks including language understanding, common sense reasoning, mathematical reasoning, summarization, coding and ability to generate truthful answers.

Packaged as an NVIDIA NIM microservice, the model is optimized for low latency, which means faster responses for users, and high throughput, which corresponds to higher computational efficiency in production.

In some cases, developers may want an even smaller version of the model to run on a smartphone or an embedded device like a robot. To do so, they can download the 8-billion-parameter model and, using NVIDIA AI Foundry, prune and distill it into a smaller, optimized neural network customized for enterprise-specific applications.

The AI Foundry platform and service offers developers a full-stack solution for creating a customized foundation model packaged as a NIM microservice. It includes popular foundation models, the NVIDIA NeMo platform and dedicated capacity on NVIDIA DGX Cloud. Developers using NVIDIA AI Foundry can also access NVIDIA AI Enterprise, a software platform that provides security, stability and support for production deployments.

Since the original Mistral-NeMo-Minitron 8B model starts with a baseline of state-of-the-art accuracy, versions downsized using AI Foundry would still offer users high accuracy with a fraction of the training data and compute infrastructure.

Harnessing the Perks of Pruning and Distillation To achieve high accuracy with a smaller model, the team used a process that combines pruning and distillation. Pruning downsizes a neural network by removing model weights that contribute the least to accuracy. During distillation, the team retrained this pruned model on a small dataset to significantly boost accuracy, which had decreased through the pruning process.

The end result is a smaller, more efficient model with the predictive accuracy of its larger counterpart.

This technique means that a fraction of the original dataset is required to train each additional model within a family of related models, saving up to 40x the compute cost when pruning and distilling a larger model compared to training a smaller model from scratch.

Read the NVIDIA Technical Blog and a technical report for details.

NVIDIA also announced this week Nemotron-Mini-4B-Instruct, another small language model optimized for low memory usage and faster response times on NVIDIA GeForce RTX AI PCs and laptops. The model is available as an NVIDIA NIM microservice for cloud and on-device deployment and is part of NVIDIA ACE, a suite of digital human technologies that provide speech, intelligence and animation powered by generative AI.

Experience both models as NIM microservices from a browser or an API at ai.nvidia.com.

See notice regarding software product information.

LINK:	https://blogs.nvidia.com/blog/mistral-nemo-minitron-8b-small-language-...
	See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

23/06/2026

PBS Selects LTN to Power Nationwide IP Video Network

Share Copy link Facebook X Linkedin Bluesky Email...

23/06/2026

PBS selects LTN for nationwide IP video network

LTN, a global leader in IP-based video transport and network services, today announced that PBS has selected LTN as its IP video partner to modernize and future...

23/06/2026

The LiveU Q Era Arrives in ANZ with the LU900Q at ABE2026

LiveU will introduce its Q Era to Australia and New Zealand for the first time at ABE2026 on Stand No. 25, (July 30 31). Leading the showcase is the LU900Q, a n...

23/06/2026

Miri Technologies Ships V410 Live 4K Video Encoder-Decode...

Miri Technologies Inc. has begun shipping its highly anticipated V410 live 4K video encoder/decoder for streaming, IP-based production workflows and AV-over-IP ...

23/06/2026

DHD SX2 and TX2 Consoles Go On-Air at Radio Tzafon

DHD audio reports the completion of an upgrade to the audio production facilities at the Galilee headquarters of Radio Tzafon. The station broadcasts two progra...

23/06/2026

Nagravision Launches Nagra Venturi Security Offering

Share Copy link Facebook X Linkedin Bluesky Email...

23/06/2026

ITN Expands Programmatic Local TV Platform

Share Copy link Facebook X Linkedin Bluesky Email...

23/06/2026

Warner Bros. Discovery Taps AWS for New AI-Powered Ad Tech

Share Copy link Facebook X Linkedin Bluesky Email...

23/06/2026

Study: Younger Viewers More Distracted But More Receptive to Ads

Share Copy link Facebook X Linkedin Bluesky Email...

23/06/2026

Chilevisin, ClaroVTR Tap Pixop for 4K FIFA World Cup Feed

Share Copy link Facebook X Linkedin Bluesky Email...

23/06/2026

Imagine Communications Names Greg Garmon as Senior Vice P...

Multifaceted Growth Executive Brings 20+ Years of Experience Leading Organizations Across Tech and M&E Imagine Communications today announced the appointment ...

23/06/2026

Australians in Film and Screen Australia's talent development initiative UNTAPPED returns for 2026

Australians in Film and Screen Australia's talent development initiative UNT...

23/06/2026

Visual Productions Unveils RdmRelay2 Four-channel Relay Control at InfoComm 2026

Visual Productions Unveils RdmRelay2 Four-channel Relay Control at InfoComm 2026 Brie Clayton June 22, 2026 0 Comments New Relay Solution Combines DMX, ...

23/06/2026

SMPTE Makes Its Standards Freely Accessible, Opening Standards Library to the Global Media Technology Community

SMPTE Makes Its Standards Freely Accessible, Opening Standards Library to the Gl...

23/06/2026

RT AND COMMUNITY FOUNDATION IRELAND ANNOUNCE RT TOY SHOW APPEAL GRANT AWARDS 2026

The RT Toy Show Appeal has raised over 31 million since its inception in 2020 ...

23/06/2026

NVIDIA Powers Over 400 of the World's 500 Fastest Supercomputers

News Highlights: NVIDIA technology runs 81% of the TOP500 and 90% of the systems new to the list. 26 systems on the TOP500 adopted the NVIDIA Grace CPU, up ei...

23/06/2026

How Businesses Are Building Specialized AI They Can Trust

Companies are asking how to build specialized AI that fits with the way their workflows actually run. The first wave of enterprise AI was about access. Compan...

23/06/2026

June 22, 2026

Newly identified molecule strengthens the eye's response to damage in retinal disease Scripps Research discovery finds that restoring the naturally occurrin...

22/06/2026

Behind the Mic: SportsCenters Lisa Cohn to Retire This June From ESPN as Longest-Tenured Anchor

Behind The Mic provides a roundup of recent news regarding on-air talent, includ...

22/06/2026

Cosm Appoints David Ho as Chief Legal Officer

Cosm has announced the appointment of David Ho as Chief Legal Officer, a newly created executive role reporting to President and CEO Jeb Terry. Ho will oversee ...

22/06/2026

Warner Bros. Discovery and AWS Announce AI-Powered Advertising Technology Platform

Warner Bros. Discovery and Amazon Web Services (AWS) have announced the developm...

22/06/2026

Daktronics Completes Audio Control System Upgrade at Petco Park for San Diego Padres

Daktronics has completed an audio control system upgrade at Petco Park in San Di...

22/06/2026

Accelerate Media Names John Willi President, Launches Accelerate Sports Network

Accelerate Media has named John Willi as President and announced the launch of the Accelerate Sports Network (ASN), a prep sports media and streaming platform c...

22/06/2026

AWSN to Air 3XBA Womens Basketball Tournament Live June 26-27

All Women's Sports Network (AWSN) and 3XBA (3 3 Basketball Association) have announced live television coverage of the annual 3XBA tournament on Friday, Jun...

22/06/2026

OWL AI Appoints Jay Prasad as Chief Executive Officer

OWL AI has announced the appointment of Jay Prasad as Chief Executive Officer and member of the Board of Directors. Prasad succeeds Josh Gwyther, who has served...

22/06/2026

CP Communications Provides RF Support for Inside the NBA at 2026 NBA Finals

CP Communications delivered RF video and audio support for TNT's Inside the NBA at the 2026 NBA Finals, providing main show coverage in San Antonio and ea...

22/06/2026

Polymarket and GRID Partner to Integrate Esports Data and Streaming into Trading Platform

Polymarket has announced a partnership with GRID, an official esports data platf...

22/06/2026

SVG New Sponsor Spotlight: Metinteractive's Rachel Mele, Ken Cyr on Building Technology Backbones for Sports Venues

As sports venues continue to evolve into more video-centric, fan-engagement-driv...

22/06/2026

SVG All-Stars: Corbin Perkins, Chief Engineer, Victory+

As the regional sports production scene shifts toward streaming, this Texan helps lead the engineering behind Victory+'s growing live platform...

22/06/2026

Meet the 2026 Sundance Institute Documentary Edit Intensive Fellows

By Kristin Feeley, Director, Documentary Film & Artist Programs the memories of your elders [are] a scaffolding for you to build your identity on - and t...

22/06/2026

Blade joins CEDAR Audio Icons line-up

New hyper-resolution analyser EQ revealed CEDAR Audio's all-new Icons plug-in series has just gained its newest member, Blade. Described by the compan...

22/06/2026

Sampleson release Aeronaut

Turn any live input into a cinematic soundscape Designed for use in the studio and on stage, Sampleson's latest creation is capable of taking any audio ...

22/06/2026

ADDAC System's new Four Strings Series

Adds guitar strings to Eurorack rigs ADDAC System are renowned for their weird and wonderful synth designs, and their line-up includes plenty of gear that&#...

22/06/2026

FIFA World Cup 2026 fever grows, as more than one third of Australians tune in to SBS coverage

FIFA World Cup 2026 fever grows, as more than one third of Australians tune in ...

22/06/2026

NAGRA Venturi - Turning Piracy Intelligence into Measurable Business Impact

In our latest blog, Tim Pearson explores NAGRA Venturi, the new streaming security solution for the AI era from NAGRAVISION. Designed to aggregate and analyze ...

22/06/2026

Xumo Expands Contextual Targeting Capabilities Through Gracenote and IRIS.TV Integrations

Expanded integrations give advertisers access to distinct contextual signals acr...

22/06/2026

Greg Garmon Joins Imagine as Senior VP, Americas Video Sales

Share Copy link Facebook X Linkedin Bluesky Email...

22/06/2026

Kaleidescape Breaks the 8K and 4:4:4 Barriers

Share Copy link Facebook X Linkedin Bluesky Email...

22/06/2026

Xilica introduces Dynamic Voice Lift in new Designer

Xilica today announced the release of Dynamic Voice Lift, a new feature in Xilica Designer v4.12 that brings adaptive speech reinforcement to large meeting spac...

22/06/2026

NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations

Telecom operators have seen remarkable returns from using generative AI to automate network management, customer care and back-office operations. Most of that i...

22/06/2026

Official trailer released for Katie Price: Nothing to Hide, coming to Sky and NOW on 8 July

Monday 22 June 2026 Official trailer released for Katie Price: Nothing to Hide,...

22/06/2026

Eco Wave Power Turns Waves Into Watts With NVIDIA AI Infrastructure and Digital Twins

The next era of AI will not be defined by compute alone. Its growth will be dete...

22/06/2026

NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory

Mission, Vision and Veritas - new Los Alamos National Laboratory (LANL) supercom...

22/06/2026

From Materials Simulation to Experimental Astronomy, New NVIDIA AI Software Unlocks Scientific Discoveries

At the ISC conference running in Hamburg this week, NVIDIA is introducing new so...

22/06/2026

NAIRR Science Program Reshapes Scientific Research, Powered by NVIDIA AI Infrastructure

For the past two years, the U.S. National Science Foundation's National Arti...

22/06/2026

At ISC, JUPITER Shows What Exascale Science Looks Like

JUPITER, Europe's first exascale supercomputer at Germany's Forschungszentrum J lich, runs on NVIDIA Grace Hopper Superchips and NVIDIA Quantum-X800 Inf...

View most recent headlines