Sony Pixel Power calrec Sony

Mixture of Experts Powers the Most Intelligent Frontier AI Models, Runs 10x Faster on NVIDIA Blackwell NVL72

03/12/2025

The top 10 most intelligent open-source models all use a mixture-of-experts architecture.

Kimi K2 Thinking, DeepSeek-R1, Mistral Large 3 and others run 10x faster to enable one-tenth the cost per token on NVIDIA GB200 NVL72.

A look under the hood of virtually any frontier model today will reveal a mixture-of-experts (MoE) model architecture that mimics the efficiency of the human brain.

Just as the brain activates specific regions based on the task, MoE models divide work among specialized experts, activating only the relevant ones for every AI token. This results in faster, more efficient token generation without a proportional increase in compute.

The industry has already recognized this advantage. On the independent Artificial Analysis (AA) leaderboard, the top 10 most intelligent open-source models use an MoE architecture, including DeepSeek AI's DeepSeek-R1, Moonshot AI's Kimi K2 Thinking, OpenAI's gpt-oss-120B and Mistral AI's Mistral Large 3.

However, scaling MoE models in production while delivering high performance and low cost per token is notoriously difficult. The extreme codesign of NVIDIA GB200 NVL72 systems combines hardware and software optimizations for maximum performance and efficiency, making it practical and straightforward to scale MoE models.

The Kimi K2 Thinking MoE model - ranked as the most intelligent open-source model on the AA leaderboard - sees a 10x performance leap on the NVIDIA GB200 NVL72 rack-scale system compared with NVIDIA HGX H200. This 10x performance leap, also seen on other MoE models such as DeepSeek-R1 and Mistral Large 3, enables one-tenth the cost per token and underscores why NVIDIA's full-stack inference platform is the key to unlocking their full potential.

What Is MoE, and Why Has It Become the Standard for Frontier Models? Until recently, the industry standard for building smarter AI was simply building bigger, dense models that use all of their model parameters - often hundreds of billions for today's most capable models - to generate every token. While powerful, this approach requires immense computing power and energy, making it challenging to scale.

Much like the human brain relies on specific regions to handle different cognitive tasks - whether processing language, recognizing objects or solving a math problem - MoE models comprise several specialized experts. For any given token, only the most relevant ones are activated by a router. This design means that even though the overall model may contain hundreds of billions of parameters, generating a token involves using only a small subset - often just tens of billions.

Like the human brain uses specific regions for different tasks, mixture-of-experts models use a router to select only the most relevant experts to generate every token. By selectively engaging only the experts that matter most, MoE models achieve higher intelligence and adaptability without a matching rise in computational cost. This makes them the foundation for efficient AI systems optimized for performance per dollar and per watt - generating significantly more intelligence for every unit of energy and capital invested.

Given these advantages, it is no surprise that MoE has rapidly become the architecture of choice for frontier models, adopted by over 60% of open-source AI model releases this year. Since early 2023, it's enabled a nearly 70x increase in model intelligence - pushing the limits of AI capability.

Since early 2025, nearly all leading frontier models use MoE designs. Our pioneering work with OSS mixture-of-experts architecture, starting with Mixtral 8x7B two years ago, ensures advanced intelligence is both accessible and sustainable for a broad range of applications, said Guillaume Lample, cofounder and chief scientist at Mistral AI. Mistral Large 3's MoE architecture enables us to scale AI systems to greater performance and efficiency while dramatically lowering energy and compute demands.

Overcoming MoE Scaling Bottlenecks With Extreme Codesign Frontier MoE models are simply too large and complex to be deployed on a single GPU. To run them, experts must be distributed across multiple GPUs, a technique called expert parallelism. Even on powerful platforms such as the NVIDIA H200, deploying MoE models involves bottlenecks such as:

Memory limitations: For each token, GPUs must dynamically load the selected experts' parameters from high-bandwidth memory, causing frequent heavy pressure on memory bandwidth.

Latency: Experts must execute a near-instantaneous all-to-all communication pattern to exchange information and form a final, complete answer. However, on H200, spreading experts across more than eight GPUs requires them to communicate over higher-latency scale-out networking, limiting the benefits of expert parallelism.

The solution: extreme codesign.

NVIDIA GB200 NVL72 is a rack-scale system with 72 NVIDIA Blackwell GPUs working together as if they were one, delivering 1.4 exaflops of AI performance and 30TB of fast shared memory. The 72 GPUs are connected using NVLink Switch into a single, massive NVLink interconnect fabric, which allows every GPU to communicate with each other with 130 TB/s of NVLink connectivity.

MoE models can tap into this design to scale expert parallelism far beyond previous limits - distributing the experts across a much larger set of up to 72 GPUs.

This architectural approach directly resolves MoE scaling bottlenecks by:

Reducing the number of experts per GPU: Distributing experts across up to 72 GPUs reduces the number of experts per GPU, minimizing parameter-loading pressure on each GPU's high-bandwidth memory. Fewer experts per GPU also frees up memory space, allowing each GPU to serve more concurrent users and support longer input lengths.

Accelerating expert communication: Experts spread across GPUs can com
LINK: https://blogs.nvidia.com/blog/mixture-of-experts-frontier-models/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

01/04/2026

DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION

January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION Douyin Users Can Now Create And Share Videos With Stun...

04/03/2026

NBC Sports, USA Sports Extend Rights Deal with PGA

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

Broadcasters Gather in DC for NAB State Leadership Conference

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

Transforming Africa's Future Farmers: Satellite-Enabled IoT Powers Data-Driven Agribusinesses

Luxembourg, March 3, 2026 - SES, a leading space solutions company, along with I...

04/03/2026

March 02, 2026

Scripps Research welcomes healthcare innovator Joe Kiani to the Board of Directors Kiani brings decades of experience in patient safety and public service. Mar...

04/03/2026

March 03, 2026

Nanoparticle vaccine approach takes on a new target: Hepatitis C virus Scripps Research scientists reengineer critical proteins on the surface of HCV, paving th...

03/03/2026

LIV Golf, Beyond Sports Elevate Online Gaming Ecosystem with Launch of LIV Golf Fantasy and LIV X

Beyond Sports, a Sony group company, and LIV Golf, the world's golf league, ...

03/03/2026

Ilitch Sports + Entertainment Announces Launch of Detroit SportsNet

Ilitch Sports + Entertainment announces the launch of Detroit SportsNet (DSN), a year-round broadcast home for two of Detroit's franchises. With flexible op...

03/03/2026

Advanced Systems Group Promotes Gretchen Taipale to Vice President, Managed Services

Advanced Systems Group, LLC (ASG), a technology and services provider for media ...

03/03/2026

PGA of America, NBC Sports, and USA Sports Extend Media Rights Agreement Through 2033

The PGA of America, NBC Sports and USA Sports extend their media rights agreemen...

03/03/2026

HONOR, ARRI Announce Technical Collaboration to Bring ARRI Image Science into Next-Gen Consumer Devices

AI device ecosystem company HONOR enters into a strategic technical collaboratio...

03/03/2026

Telos Alliance Partners with College Radio Foundation to Support College Broadcasters

Cleveland's Telos Alliance, pioneers in broadcast technology for 30 years, l...

03/03/2026

Sennheiser Relaunches MD 9235 Wireless Mic Head

The MD 9235 microphone head for wireless handhelds has been a firm favorite with many engineers and artists for its ability to cut through high on-stage levels ...

03/03/2026

Haivision to Showcase Private 5G and Live Video Contribution Innovations at MWC 2026

Haivision Systems Inc. (Haivision), a global provider of mission-critical, real-...

03/03/2026

BMG Expands Washington Broadcast Center with 3 New TV Studios and Podcast Studio for Media Clients

Broadcast Management Group (BMG) announces the expansion of its 62,000-square-fo...

03/03/2026

Closing the Loop: Maroon 5 and the End of the Analog Era

Maroon 5's musical tour in 2025 marked a leap forward in live audio as Monitor Engineer Dave Rupsch utilized Sennheiser's all-digital Spectera wireless ...

03/03/2026

SVG in Indy: Pacers Sports & Entertainment Finds Production Sweet Spot in ST 2110-Based Control Center

Designed specifically for pro basketball, the renovated space at Gainbridge Fiel...

03/03/2026

Lawo Appoints Jamie Dunn CEO

As part of the move, former CEO Phillipp Lawo joins the broadcast-tech provider's Supervisory Board Lawo has announced appointment of Jamie Dunn as chief e...

03/03/2026

NBC Turns Back the Clock to 1990s for NBA Coast 2 Coast' Tuesday

A team of legendary announcers and analysts and a classic graphics look will bring the past to life NBC Sports and Peacock will return to yesteryear for tonigh...

03/03/2026

Sundance Film Festival: CDMX 2026 Returns for Its Third Edition

From April 30 to May 3, Sundance Film Festival: CDMX 2026 will offer a selection of exciting independent cinema. Mexico City, March 3, 2026 - At a moment of he...

03/03/2026

How Multi-Format Readers' Are Redefining Reading in the UK's National Year of Reading

For many, finding time or headspace to pick up a book can feel out of reach, but...

03/03/2026

Rohde & Schwarz and Realtek demonstrate first test solution for Bluetooth LE High Data Throughput (HDT)

Rohde & Schwarz and Realtek demonstrate first test solution for Bluetooth LE Hi...

03/03/2026

Sediba Scriptwriting Training Programme - Matatiele (Eastern Cape)

The National Film and Video Foundation (NFVF) invites aspiring and emerging filmmakers from Matatiele and surrounding areas to apply for the Sediba Scriptwritin...

03/03/2026

Clear-Com Supplies Cloud-based Communications System for SaxaVord Spaceport

eds3_5_jq(document).ready(function($) { $(#eds_sliderM519).chameleonSlider_2_1({ content_source:......

03/03/2026

Magellan AI Integrates Nielsen DMA Data to Bring Local Market Measurement to Podcast Attribution

Nielsen's DMA data gives Magellan AI users a standardized way to measure th...

03/03/2026

Lawo Promotes Jamie Dunn to CEO

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Elements To Showcase Newly Unveiled GRID NAS Platform At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Moments Lab To Feature Agentic AI For Video Workflows At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Marshall Electronics Launches Compact CV356-10X Full HD C...

Marshall Electronics premieres the CV356-10X, its latest compact 10X camera that offers Full HD with simultaneous SDI and HDMI outputs, at NAB 2026 (Booth C8339...

03/03/2026

farmerswife and Cirkus to Showcase Smarter Media Workflow...

farmerswife, the industry-leading enterprise operations platform for broadcast and post-production, today announced it will exhibit at NAB Show 2026 in Las Vega...

03/03/2026

Manfrotto ONE Hybrid Tripod Wins iF Design Award 2026

Manfrotto has announced that the Manfrotto ONE Hybrid tripod has won the iF DESIGN AWARD 2026, one of the world's most respected design honours. Selected ...

03/03/2026

DHD to Introduce Latest Generation Broadcast Audio Mixers...

DHD is expanding the capabilities of its DX2, RX2, SX2 and TX2 broadcast audio mixers, RM1 portable production unit and XC3/XD3/XS2 processing cores with the in...

03/03/2026

Synamedia and MoMe launch first streaming CDN in Spain

Leading video software provider Synamedia and MoMe, a leading Spanish consultancy and systems integrator, today announced the launch of Spain's first stream...

03/03/2026

Long-Awaited ATSC 3.0 Rulemaking Overshadows NAB Show Expectations

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Audio Tech at NAB Show: Are We in the Second Wave' of IP?

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

IP's Impact on Imaging Tech on Full Display at NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Live Production Over IP in 2026: Software-Defined Everything

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

NAB Show Leverages Revitalized LVCC To Reflect M&E Transformation

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Home Post Production Strengthens Factual and Natural Hist...

Home Post Production has further expanded its factual, unscripted, and entertainment capabilities with the acquisition of Picture Shop Bristol, a leading post h...

03/03/2026

Full Year 2025 Results

Luxembourg, 2 March 2026 -- SES S.A. fully consolidates Intelsat from 17 July 2025 and announces financial results for the year ended 31 December 2025 FY25 Pe...

03/03/2026

Iyuno Taps Dante AV to Sync Audio and Video Content

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

HBO Max and Paramount+ Streamers to Merge

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Survey: 70% of CTV Advertisers Plan to Boost Spending in 2026

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

XGN Global, X1 Mobile Show New 5G Broadcast Smartphone

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Scripps Completes Sale of WFTX to Sun Broadcasting

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

HbbTV Association Formally Integrates DRM into Core Specification

Share Copy link Facebook X Linkedin Bluesky Email...