Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

02/06/2026

Case Study: Tennis Channel Transitions From Satellite to IP Distribution With LTN

Tennis Channel has completed a transition from satellite-based distribution to a...

02/06/2026

Daktronics Announces 2026 High School Video Summit for June 23-24

Daktronics has announced the 2026 High School Video Summit, a two-day educational event for high school educators and student production teams, taking place Jun...

02/06/2026

Fandango to Screen Telemundos FIFA World Cup 2026 Coverage at 160 Theaters Nationwide

Fandango will bring Telemundo's live Spanish-language coverage of the FIFA W...

02/06/2026

AWSN Announces Broadcast Schedule for Inaugural Professional Softball League Season

All Women's Sports Network (AWSN) has announced the live television schedule...

02/06/2026

Spalk Announces Partnerships with Ligue 1, Euroleague, and LNB

Spalk, a cloud-based multilingual commentary and production platform, has announced three new partnerships: Ligue 1 (English and Portuguese highlights), Eurolea...

02/06/2026

NHL Network Announces Stanley Cup Final Programming Schedule

NHL Network's 2026 Stanley Cup Final coverage began June 1 with NHL Tonight: Stanley Cup Final Media Day from the Carolina Hurricanes' Lenovo Center at ...

02/06/2026

Matrox Video Launches Maevex MGX Series IPMX-Ready AV-over-IP Encoders and Decoders

Matrox Video has announced the Maevex MGX Series, a lineup of IPMX-ready video e...

02/06/2026

Marshall Electronics to Exhibit IP Camera Lineup at InfoComm 2026

Marshall Electronics will exhibit at InfoComm 2026 (Booth C7521), showcasing a lineup of 4K and HD compact POV cameras for corporate, education, hospitality, wo...

02/06/2026

PAMA and Shure Extend Deadline for Mark Brunner Professional Audio Scholarship to June 15

The Professional Audio Manufacturers Alliance (PAMA) and Shure Incorporated have...

02/06/2026

MultiDyne To Debut FiberSaver-10G and New VersaBrix Frame Options at InfoComm 2026

MultiDyne Video and Fiber Optic Systems will exhibit at InfoComm 2026 (Booth C50...

02/06/2026

Advanced Systems Group Promotes Joe Marchitto to Western Regional CTO

Advanced Systems Group (ASG) has announced the promotion of Joe Marchitto to Western Regional CTO. In his new role, Marchitto will oversee system design across ...

02/06/2026

Brothers Osborne To Perform Free Concert Outside Lenovo Center Before Stanley Cup Final Game 1

The NHL has announced that Brothers Osborne will headline a free outdoor concert...

02/06/2026

CP Communications Names First Female Executives in Company History

CP Communications has announced the appointment of its first two female executives in the company's 40-year history. Tabitha Coleman has been named Vice Pre...

02/06/2026

Lighting Design Group Designs Broadcast Studio for Yahoo at 770 Broadway

The Lighting Design Group (LDG) has announced the completion of Studio C at Yahoo's headquarters at 770 Broadway in Manhattan. The studio launched April 24 ...

02/06/2026

Zee Entertainment Acquires FIFA Media Rights for India Through 2034

Zee Entertainment Enterprises Ltd. (Z) has announced a partnership with FIFA to broadcast FIFA World Cup 2026, FIFA World Cup 2030, FIFA Women's World Cup 2...

02/06/2026

Behind The Mic: Prime Sports UK Shares On Air Roster for NBA Finals; and More

Behind The Mic provides a roundup of recent news regarding on-air talent, including new deals, departures, and assignments compiled from press releases and repo...

02/06/2026

Auburn Universitys Parker Leppien on the Nonstop Growth of War Eagle Productions

The department is broadcasting NCAA Baseball Super Regionals at Plainsman Park this weekend Broadcast and production crews at Division I institutions are busy ...

02/06/2026

With Stanley Cup Final Set To Begin, NHL and Honeywell Lay Foundation for Energy-Efficient Venues

The initiative is extends beyond the league's arenas to team training rinks,...

02/06/2026

Release Rundown: What to Watch in June, From Leviticus to The Invite

Olivia Wilde, Seth Rogen, Pen lope Cruz, and Edward Norton appear in The Invite by Olivia Wilde, an official selection of the 2026 Sundance Film Festival. (Co...

02/06/2026

Cobalt Digital to Showcase End-to-End IPMX Ecosystem at InfoComm 2026, Making ST 2110 Easy for Pro AV

blueCORE standalone processors headline solutions designed to simplify the trans...

02/06/2026

Marketing Architects Expands Relationship With Nielsen To Include Integration of Media Data Engine on a National Level

The TV agency was one of the earliest adopters of Nielsen's local television...

02/06/2026

Riedel Networks Taps Gudrun Scharler as CEO

Share Copy link Facebook X Linkedin Bluesky Email...

02/06/2026

Grass Valley Enables Sky News Australia s Cloud-First New...

Grass Valley today announced that Australian News Channel (ANC), operator of Sky News Australia, has deployed Grass Valley AMPP to transform its newsroom produc...

02/06/2026

STUDIO TECHNOLOGIES INTRODUCES NEW MODEL 385 MIC INTERCOM...

Studio Technologies, a leading manufacturer of high-quality audio, video, and fiber-optic solutions, announces its new Model 385 Mic/Intercom Beltpack. The Mode...

02/06/2026

Gudrun Scharler Appointed CEO of Riedel Networks

The Riedel Group today announced the appointment of Gudrun Scharler as CEO of Riedel Networks. She succeeds Michael Martens, who has led Riedel Networks since 2...

02/06/2026

Magewell Levels-Up All-in-One Content Production with Lau...

More signals, higher quality, and outstanding ingest and streaming flexibility deliver professional results in a small, all-in-one footprint...

02/06/2026

Modena Showcases farmerswife at Mediatech 2026

farmerswife will be featured on the Modena Media & Entertainment stand at this year's Mediatech Africa 2026, giving visitors an opportunity to explore the l...

02/06/2026

PTZOptics showcases intelligent video ecosystem at InfoCo...

PTZOptics will showcase a new generation of intelligent video workflows at InfoComm 2026, June 17 19, Las Vegas. Visitors to booth N8227 will see how PTZOptics ...

02/06/2026

Roku Launches the Roku Soccer Zone

Share Copy link Facebook X Linkedin Bluesky Email...

02/06/2026

FCC Sets Deadlines for Comments in ABC License Renewals

Share Copy link Facebook X Linkedin Bluesky Email...

02/06/2026

Studio Technologies Introduces Model 385 Beltpack

Share Copy link Facebook X Linkedin Bluesky Email...

02/06/2026

Gerald Jerry Pierce, Architect of Modern Digital Cinema, Dies at 73

Share Copy link Facebook X Linkedin Bluesky Email...

02/06/2026

Discover Your Next Book-to-Screen Obsession With Our New Dedicated Hub

Back to All News Discover Your Next Book-to-Screen Obsession With Our New Dedicated Hub Product 02 June 2026 Global Link copied to clipboard Dear most vor...

02/06/2026

New LinkedIn Data: 90% of C-Suite Leaders are Continuously ...

New LinkedIn Data: 90% of C-Suite Leaders are Continuously Building Skills to Lead in the AI Era Published on Jun 2, 2026 Categories: Data and insights Lin...

01/06/2026

CBS Sports UEFA Champions League Today Studio Show Heads to Budapest for Final as Transcontinental Popularity Grows

In its sixth year, the broadcaster's coverage has become a global brand and ...

01/06/2026

AudioShake Launches End-to-End Copyright Compliance System for Mixed-Media Audio

Designed to solve a common problem in broadcasting, the automated workflow detects, identifies, removes, and documents copyrighted music AudioShake has introdu...

01/06/2026

SVG Sit-Down: Stats Perform's Charles Kaplan on 30 Years of Opta, a Busy Summer of Soccer, What's Next

The sports-analytics company combines its data with proprietary AI to help leagu...

01/06/2026

ASG Advances Joe Marchitto to Western Regional CTO

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

Scripps Stations Go Dark on DirecTV

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

MARSHALL ELECTRONICS POWERS SEAMLESS AV EXPERIENCES WITH...

Marshall Electronics is showcasing a comprehensive lineup of next-generation POV cameras, purpose-built to power today's connected AV environments, at InfoC...

01/06/2026

Adobe Announces Concept to Vector

Adobe Announces Concept to Vector Deepa Subramaniam June 1, 2026 0 Comments One of the biggest frustrations we hear from designers is how difficult it...

01/06/2026

Vampire Feature Night Patrol Graded with DaVinci Resolve Studio

Vampire Feature Night Patrol Graded with DaVinci Resolve Studio Brie Clayton June 1, 2026 0 Comments Colorist shapes dark, gritty tone for horror thri...

01/06/2026

U.S. Broadcasters Ready for Most Complex FIFA World Cup Ever

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

Broadcasters Prepare for Nation's 250th Birthday Bash

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

Broadcasters Reveal What Makes C-Band Alternatives Right for Them

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

IAMT to Offer New Educational Sessions at InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

NewsNation Launches New Podcasting Studio and Podcasts

Share Copy link Facebook X Linkedin Bluesky Email...

01/06/2026

KAULITZ & KAULITZ Season 3

Back to All News KAULITZ & KAULITZ Season 3 Entertainment 01 June 2026 GlobalGermany Link copied to clipboard Our favorite twins are back: streaming on Ne...

01/06/2026

NVIDIA Jetson Brings Agentic AI to the Physical World

Agentic AI is getting physical. At COMPUTEX on Tuesday, NVIDIA announced NVIDIA JetPack 7.2 and NVIDIA NemoClaw support on NVIDIA Jetson. JetPack 7.2 brings a...

01/06/2026

Why Financial Institutions Are Converging on Transaction Foundation Models to Build Their Own Intelligence

Financial institutions have spent years building AI: fraud models, credit models...