Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

02/01/2026

Freezing Florida: NHL's EVP, Entertainment Bob Chesterman on Taking the Winter Classic to Miami

Freezing Florida: NHL's EVP, Entertainment Bob Chesterman on Taking the Wint...

02/01/2026

NHL Winter Classic 2026: TNT Sports Prepares for First NHL Outdoor Game in Sunshine State

NHL Winter Classic 2026: TNT Sports Prepares for First NHL Outdoor Game in Sunsh...

02/01/2026

ATSC To Showcase Latest NextGen TV Developments At CES 2026

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

02/01/2026

Pearl TV To Unveil NextGen TV Converter Box Program At CES 2026

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

01/01/2026

Latin Grammy Cultural Foundation Announces 2026 Noel Schajris Scholarship

Latin Grammy Cultural Foundation Announces 2026 Noel Schajris Scholarship The scholarship will cover tuition and housing for one Berklee College of Music stud...

01/01/2026

GeForce NOW Rings In 2026 With 14 New Games in January

New year, new games, all with RTX 5080-powered cloud energy. GeForce NOW is kicking off 2026 by looking back at an unforgettable year of wins and wildly high fr...

31/12/2025

NFL Christmas Gameday on Netflix Scores Again With the Lions-Vikings Becoming the Most-Streamed NFL Game in US History

Back to All News NFL Christmas Gameday on Netflix Scores Again With the Lions-V...

30/12/2025

As the College Football Playoff Enters the Quarterfinals, ESPN Blows Out Its MegaCast Multiplatform Playbook

As the College Football Playoff Enters the Quarterfinals, ESPN Blows Out Its Meg...

30/12/2025

SVG's Best of 2025: Original Articles

SVG's Best of 2025: Original ArticlesTake a look back at all our coverage of big-time productions, game-changing technologies, and state-of-the-art new faci...

30/12/2025

L3Harris Sets Date for Fourth Quarter 2025 Earnings Release

MELBOURNE, Fla., Dec. 30, 2025 - L3Harris Technologies (NYSE: LHX) will release its fourth quarter 2025 financial results before the market opens on Thursday, J...

30/12/2025

TV Techs Most Popular Stories of 2025

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

29/12/2025

San Francisco 49ers Strike Gold With Halftime Laser Spectacular

San Francisco 49ers Strike Gold With Halftime Laser SpectacularStunning display caps $200 million renovation of Levi's Stadium techBy Dan Daley, Audio Edito...

29/12/2025

The Cup's Around the Corner: An Inside Look at Broadcast Preparations for the 2026 FIFA World Cup With FIFA's Oscar Sanchez

The Cup's Around the Corner: An Inside Look at Broadcast Preparations for th...

29/12/2025

SVG's Best of 2025: Longform Video

SVG's Best of 2025: Longform VideoWatch the standout keynote conversations, deep dives, and panel discussions from the year for free on SVG PLAY!By Brandon ...

29/12/2025

TV Tech's Top Regulatory Stories of the Year

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

27/12/2025

TV Tech's Top Data Dumps of 2025

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

26/12/2025

TV Tech's Top Streaming Stories of 2025

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

26/12/2025

TV Tech's Top Data Points of 2025

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

25/12/2025

Make Spirits Bright With Holiday Hits on GeForce NOW

Holiday lights are twinkling, hot cocoa's on the stove and gamers are settling in for a well-earned break. Whether staying in or heading on a winter getawa...

24/12/2025

What is AI good for?

What is AI good for? Posted by MTI Film on December 24, 2025 What is AI good for? What is AI good for? It's been three years since ChatGPT first cap...

24/12/2025

AI in 2026: More Collaboration, Less Hype

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

24/12/2025

Carr Lays Out FCCs 'Key Wins in 2025'

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

24/12/2025

CES: Cineverse Unveils New Features for Cinesearch

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

24/12/2025

IES, AES Promote Graham Kirk, Brienne Willcock

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

24/12/2025

Ad Tech and CTV Experts Forecast 2026's Biggest Trends

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

24/12/2025

The Boyfriend' Season 2 Unveils Heartwarming Trailer, Key Art, and Participants' Profiles

Back to All News The Boyfriend' Season 2 Unveils Heartwarming Trailer, Key...

24/12/2025

Love, Fights, and Everything in Between: Badly in Love' Returns for Season 2

Back to All News Love, Fights, and Everything in Between: Badly in Love' Returns for Season 2 Entertainment 24 December 2025 GlobalJapan Link copied t...

24/12/2025

December 23, 2025

Scripps Research study links sleep variability with sleep apnea and hypertension How consumers' digital activity trackers could enable personalized health s...

23/12/2025

How guilas Cibaeas Dominican Winter League Games Are Locally Produced for Global Audience

How guilas Cibae as Dominican Winter League Games Are Locally Produced for Glob...

23/12/2025

CAMB.AI Enables European Athletics to Offer Multi-Language Support

CAMB.AI Enables European Athletics to Offer Multi-Language SupportPlan is to eventually offer translation into all languages spoken in EuropeBy Ken Kerschbaumer...

23/12/2025

Analysis: As Sports Media Values Trend Negative, Scarcity and Quality Are King

Analysis: As sports media values trend negative, scarcity and quality are king By Callum McCarthy, Editor-at-Large Monday, December 22, 2025 - 14:08 Print ...

23/12/2025

ESPN, Disney, and NBA Return to the Animated Altcast Fray With Second Edition of Dunk the Halls'

ESPN, Disney, and NBA Return to the Animated Altcast Fray With Second Edition of...

23/12/2025

End the Year on a High Note and Donate to the Sports Broadcasting Fund Today!

End the Year on a High Note and Donate to the Sports Broadcasting Fund Today!By Ken Kerschbaumer, Editorial Director Tuesday, December 23, 2025 - 12:25 pm P...

23/12/2025

L3Harris Receives Letter of Intent from Kratos Defense for Production of Large Hypersonic Solid Rocket Motors

A Zeus motor is hot fire tested at L3Harris' Camden, Arkansas, solid rocket ...

23/12/2025

FCC Bans All New Foreign-Made Drones

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Gray Media Renews Its NBC Affiliation Agreements

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Lightware to showcase breakthrough Google Meet and TPN MM...

Lightware will exhibit several major product innovations at ISE 2026, including the new USB-C BOOSTER-V1, Google Meet. integration for various Taurus UCX models...

23/12/2025

Nielsen, Roku Expand Measurement Partnership

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

PwC: Streaming Market Shifting to 'Scale and Sustainability'

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Inside the Gray Innovation Lab

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

ESPN Renews Deal for Heisman Trophy Coverage

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Gray Media to Acquire WBBJ from Bahakel Communications

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Taking the Stage at Carnegie Hall-On a Global Scale

Taking the Stage at Carnegie Hall-On a Global Scale Boston Conservatory Orchestra students reflect on their epic concert marking the 80th session of the UN Gene...

23/12/2025

Netflix's 'The Great Flood' and 'Culinary Class Wars 2' Top Global Charts Simultaneously

Back to All News Netflix's The Great Flood and Culinary Class Wars 2 Top Gl...

23/12/2025

'Stranger Things' By the Numbers: How the Global Phenomenon Shaped Culture

Back to All News Stranger Things By the Numbers: How the Global Phenomenon Shap...

23/12/2025

Boost Performance with a System Effectiveness Review

Experience the power of WO Automation for Radio's newest service, the System Effectiveness Review. Designed to help you achieve more, a System Effectiveness...

23/12/2025

How Steamy Can It Get? Single's Inferno' Season 5 Premieres January 20, Previews All-Out Flirting War in Sizzling Teaser

Back to All News How Steamy Can It Get? Single's Inferno' Season 5 Pre...

23/12/2025

33 Million Global Viewers on Netflix Watched Jake Paul vs. Anthony Joshua's Epic Six-Round Battle

Back to All News 33 Million Global Viewers on Netflix Watched Jake Paul vs. Ant...