Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

21/11/2025

FCC Votes to Clear at Least 100MHz of Upper C-Band Spectrum

WASHINGTON The Federal Communications Commission by a 3-0 vote adopted a Notice of Proposed Rulemaking (NPRM) to advance Congress's mandate to clear a minim...

21/11/2025

Spectrum Expands 4K Content to Apple TV 4K and Roku Devices

STAMFORD, Conn. Charter's Spectrum has expanded the devices that can offer 4K content on the Spectrum TV app to compatible Apple TV 4K and Roku devices....

21/11/2025

Study: Salaries in Content, Connectivity Industries Continue To Grow

NAPERVILLE, Ill. Media industry employers are continuing their multiyear trend of increasing salaries for all worker segments but lag general industry raises, s...

21/11/2025

NAB Opens Nominations for 2026 Technology Awards

WASHINGTON The National Association of Broadcasters said it is accepting nominations for the 2026 NAB Technology Awards, honors that recognize excellence in bro...

21/11/2025

AAT Introduces Automated RF Line Analysis

American Amplifier Technologies has released a vector network analysis module....

21/11/2025

The Best Movie Musicals on Every Streaming Platform

The Best Movie Musicals on Every Streaming Platform From Wicked to The Sound of Music, heres where to stream all the classic movie musicals and recent hits on...

20/11/2025

MLB Media-Rights Shakeup: NBC's New Three-Year Deal Covers Sunday Night Baseball,' Peacock-Exclusive Games, and More

MLB Media-Rights Shakeup: NBC's New Three-Year Deal Covers Sunday Night Bas...

20/11/2025

MLB Media-Rights Shakeup: New Deal Will Bring 30 National Games to ESPN's Linear Platform, MLB.TV Exclusively to ESPN App

MLB Media-Rights Shakeup: New Deal Will Bring 30 National Games to ESPN's Li...

20/11/2025

MLB Media-Rights Shakeup: Netflix Lands Opening Night, Home Run Derby, Field of Dreams

MLB Media-Rights Shakeup: Netflix Lands Opening Night, Home Run Derby, Field of ...

20/11/2025

MLB Media-Rights Shakeup Overview: ESPN, NBCU, Netflix Ink Three-Year Deals

MLB Media-Rights Shakeup Overview: ESPN, NBCU, Netflix Ink Three-Year DealsESPN gets new 30-game package, MLB.TV; NBC extends Sunday nights; Netflix adds tentpo...

20/11/2025

SVG Students To Watch: Henry Thuss, Indiana University

SVG Students To Watch: Henry Thuss, Indiana UniversityThe Southern California product has his goals set on the front benchBy Brandon Costa, Director of Digital ...

20/11/2025

Done+Dusted's Guy Carrington on Creating the Spectacular League of Legends World Championship Opening Ceremony

Done+Dusted's Guy Carrington on Creating the Spectacular League of Legends W...

20/11/2025

FIA Extreme H World Cup Host Broadcaster Aurora Goes Inside the Production of the Hydrogen-Fuelled Motorsport

FIA Extreme H World Cup Host Broadcaster Aurora Goes Inside the Production of th...

20/11/2025

Platinum White Paper: Amagi Utilizes Cloud Production for Sports Events - Multi-Camera Live Workflow with Remote Commentary and Graphics

Platinum White Paper: Amagi Utilizes Cloud Production for Sports Events - Multi-...

20/11/2025

2025 Sports Broadcasting Hall of Fame: Marc Herklotz, Steady Hand Behind the Scenes

2025 Sports Broadcasting Hall of Fame: Marc Herklotz, Steady Hand Behind the Sce...

20/11/2025

NFL Deep Dive: How 32 Cameras at Each Stadium Drive Virtual Measurement, Boundary Replays, and Skeletal Tracking

NFL Deep Dive: How 32 Cameras at Each Stadium Drive Virtual Measurement, Boundar...

20/11/2025

Zodiac Killer Project Refuses to Kill Its Darlings With an Avant-Garde Essay Film

Charlie Shackleton attends the 2025 Sundance Film Festival premiere of Zodiac K...

20/11/2025

Cracking the Code for Resilient, Secure Space-based Communications

L3Harris has achieved NSA Cybersecurity Directorate certification for its KSV-650 space hub end cryptographic unit, ensuring secure, adaptable communications fo...

20/11/2025

Continued Commitment and Investment in Poland

L3Harris is strengthening its long-term partnership with Poland through the opening of a new facility in Warsaw to serve as a central business hub for regional ...

20/11/2025

Oklahoma Community Television Overcomes ATSC 3.0 Translator Challenge

CINCINNATI Oklahoma Community Television (OCT) has selected GatesAir and Triveni Digital as key technology partners in an ATSC 3.0 deployment that establishes a...

20/11/2025

FCC Launches Wide-Ranging Examination of Network Affiliate Relations

WASHINGTON The Federal Communications Commission has opened a wide-ranging inquiry into the relations between broadcast networks and their affiliates that could...

20/11/2025

IAB: Creator Economy Ad Spend Now Dwarfs Ad Spend for Total Media Industry

NEW YORK Ad spend in the creator economy has more than doubled since 2021 from $13.9B to $29.5B in 2024 and that amount is projected to reach $37 billion in 202...

20/11/2025

Graduate Spotlight: Krysta Mirsik DePuy

Graduate Spotlight: Krysta Mirsik DePuy The educator from Hackettstown, New Jersey, shares how it took her 11 years to find the right graduate program for her...

20/11/2025

Rise Awards 2025 Celebrate the Industry's Best

LONDON The winners of the Rise Awards 2025, which recognize women and companies whose achievements have stood out in the media technology industry, have been an...

20/11/2025

Lionsgate, Debmar-Mercury Launch MovieSphere Gold Diginet

SANTA MONICA & NEW YORK Lionsgate's Worldwide Television Distribution Group and Debmar-Mercury have launched MovieSphere Gold in more than 30 million homes....

20/11/2025

Clear-Com Unveils New 4-Channel HelixNet Beltpack

ALAMEDA, Calif. Clear-Com has introduced its four-channel HelixNet beltpack, a next-generation advancement of its widely used two-channel model, and anticipates...

20/11/2025

Kiswe Core Aims to Simplify Cloud-Based Streaming Distribution

NEW PROVIDENCE, N.J. Streaming technology and services provider Kiswe has launched Kiswe Core, a next-generation cloud-based platform designed to simplify the s...

20/11/2025

Dolby Partners Up With Bay Area Super Bowl Host Committee

MOUNTAIN VIEW, Calif. Dolby Laboratories has signed on as an official signature partner of the Bay Area Host Committee (BAHC) for the 2026 Super Bowl and the we...

20/11/2025

Telus TV+ Expands Distribution to Samsung and LG Smart TVs

STUTTGART, Germany 3 Screen Solutions (3SS) has announced that the Canadian multiservice operator Telus has launched its Telus TV+ entertainment platform on Sam...

20/11/2025

MLB Strikes Rights Deals With ESPN, NBCUniversal, Netflix

Major League Baseball has announced deals with ESPN, NBCUniversal and Netflix covering 2026-2028 that make ESPN the exclusive distributor for MLB.TV, put regula...

20/11/2025

Berklee Presents Holiday Extravaganza: Duke Ellingtons The Nutcracker Suite and GenNext All-Stars' Noel

Berklee Presents Holiday Extravaganza: Duke Ellingtons The Nutcracker Suite and ...

20/11/2025

Netflix Celebrates the Launch of 'Troll 2' With a Spectacular Show in the Sky

Back to All News Netflix Celebrates the Launch of Troll 2 With a Spectacular Sh...

20/11/2025

Into the Omniverse: How Smart City AI Agents Transform Urban Operations

Editor's note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners and enterprises can transform their workflows u...

20/11/2025

Ultimate Cloud Gaming Is Everywhere With GeForce NOW

The NVIDIA Blackwell RTX upgrade is nearing the finish line, letting GeForce NOW Ultimate members across the globe experience true next-generation cloud gaming ...

20/11/2025

The Largest Digital Zoo: Biology Model Trained on NVIDIA GPUs Identifies Over a Million Species

Tanya Berger-Wolf's first computational biology project started as a bet wit...

19/11/2025

Sphere Sound Kicks Up Its Heels at Radio City Music Hall

Sphere Sound Kicks Up Its Heels at Radio City Music HallUltra-immersive system from the singular Las Vegas venue is also ready for sportsBy Dan Daley, Audio Edi...

19/11/2025

The Future of Stadium and Arena Video Control Rooms: Tech Leaders Talk 4K, IP, and More

The Future of Stadium and Arena Video Control Rooms: Tech Leaders Talk 4K, IP, a...

19/11/2025

One Year at Intuit Dome: Walking Through the Stunning Arena's Advanced Video Production Layout with the LA Clippers

One Year at Intuit Dome: Walking Through the Stunning Arena's Advanced Video...

19/11/2025

SVG Sit-Down: IMG's Francois Westcombe Goes Inside IMG's Annual Digital Trends Report

SVG Sit-Down: IMG's Francois Westcombe Goes Inside IMG's Annual Digital ...

19/11/2025

Inside Brighton & Hove Albion's Fan-First Media Machine

Inside Brighton & Hove Albion's fan-first media machine By George Bevir Tuesday, November 18, 2025 - 10:10 Print This Story Brighton & Hove Albion'...

19/11/2025

Rock Chalk Rebuild, Part 1: Kansas Athletics Brings New Production Power to Rebuilt David Booth Kansas Memorial Stadium

Rock Chalk Rebuild, Part 1: Kansas Athletics Brings New Production Power to Rebu...

19/11/2025

Rock Chalk Rebuild, Part 2: A Pro's Guide to Kansas Athletics' New Video-Production Infrastructure

Rock Chalk Rebuild, Part 2: A Pro's Guide to Kansas Athletics' New Video...

19/11/2025

MLB Media Rights Shakeup: NBC Inks Three-Year Deal for Sunday Night Baseball, Peacock Exclusive Game, and More

MLB Media Rights Shakeup: NBC Inks Three-Year Deal for Sunday Night Baseball, Pe...

19/11/2025

Sabar Bonda (Cactus Pears) Takes a Tender View on Queer Relationships

Rohan Parashuram Kanawade attends the premiere of Sabar Bonda (Cactus Pears) at the 2025 Sundance Film Festival at Egyptian Theatre on January 26, 2025, in Pa...

19/11/2025

Sundance Institute-Supported Seeds, Apocalypse in the Tropics, Selena y Los Dinos, and More Earn IDA Documentary Awards Nominations

Awards season is heating up, and the International Documentary Association just ...

19/11/2025

L3Harris Breaks Ground on Arkansas Advanced Propulsion Facilities

Aerojet Rocketdyne President Ken Bedingfield and Arkansas Gov. Sarah Huckabee Sanders were joined by state and local officials to celebrate the groundbreaking o...

19/11/2025

L3Harris and PentenAmio Team to Advance Sovereign Cryptographic Capability and Innovation in the UK, Australia and Canada

L3Harris and PentenAmio formalise their teaming agreement at MilCIS 2025, streng...

19/11/2025

The Gauge: Poland | October 2025

October brings a larger audience in front of TV screens. The cooler autumn weather and the continuation of new programming schedules brought increases in both t...

19/11/2025

Nexstar Urges FCC to Waive Ownership Rules, Quickly Approve Tegna Acquisition

IRVING, Texas Nexstar Media Group and Tegna filed applications on Nov. 18 with the Federal Communications Commission (FCC) seeking its consent to transfer broad...

19/11/2025

Vitec Buys Datapath, Expanding Product Portfolio, Engineering, Support

Vitec has acquired Datapath, a developer of real-time video processing for large-scale video walls, AVoIP content distribution and KVM control in mission-critic...