Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

25/06/2026

Launching a Career in Broadcast Engineering: Academic Paths and Essential Certifications

Launching a Career in Broadcast Engineering: Academic Paths and Essential Certif...

25/06/2026

SVG Students To Watch: Jude Kieffer, Ball State University

This superstar shooter/storyteller from Central Indiana hopes to make his mark in the blossoming sports-documentary and -features space In the live-sports-vid...

25/06/2026

Presidio and NHL Renew Multiyear North American Technology Partnership

Presidio and the National Hockey League have announced a multiyear renewal of their North American partnership. Presidio will remain an Official Technology Inno...

25/06/2026

Strike Fighter League Hits the Industry as First Professional Air Combat Sport

Strike Fighter League (SFL) is the world's first professional air combat digital sport that combines elite human performance and physical immersion with cut...

25/06/2026

Rise Reveals 2026 Worldwide Mentoring Cohorts to Support Future Industry Leaders

Rise, the award-winning advocacy group for gender diversity in the broadcast and media technology sector, is pleased to announce the global mentoring cohort for...

25/06/2026

MLB Network To Air American Association of Professional Baseball All-Star Game for First Time on July 15

The 2026 American Association of Professional Baseball (AAPB) All-Star Game will...

25/06/2026

Mediaproxy Partners with HVS for U.S. Broadcast Market

Mediaproxy has named Heartland Video Systems (HVS) as its exclusive partner for US television broadcasting. The Wisconsin-based systems integrator will represen...

25/06/2026

Backblaze Inks Five-Year Multi-Exabyte Data Storage Agreement with CoreWeave

Backblaze has formed an agreement with CoreWeave to create The Essential Cloud for AI. Under the multi-exabyte, $335 million agreement, Backblaze will provide...

25/06/2026

Clear-Com FreeSpeak Cell Tested by RTL Deutschland on 5G Network at Nrburgring

Clear-Com has announced the successful deployment and testing of FreeSpeak Cell by RTL Deutschland during a live event production at the N rburgring race circui...

25/06/2026

Mobile TV Group Launches Full-Stack MTVG Production Platform, Powers Angels Broadcast Television

Mobile TV Group (MTVG) has announced the launch of the MTVG Production Platform,...

25/06/2026

Sony Pictures Entertainment Announces $100 Million Investment in Cosm

Sony Pictures Entertainment (SPE) has announced a $100 million strategic investment in Cosm as lead investor in the company's Series C financing round, acqu...

25/06/2026

FOX Sports Renews Concacaf Gold Cup Rights and Adds Nations League Through 2029

FOX Sports and Concacaf have announced a multi-year media rights agreement making FOX Sports the U.S. English-language home of the Concacaf Gold Cup and Concaca...

25/06/2026

InfoComm 2026: Daktronics and Grass Valley Win rAVe Pubs Best Solution for Large Venue or Live Events

Daktronics and Grass Valley have received the rAVe Pubs Best Solution for Large ...

25/06/2026

Sports and Dramas Drive April Viewing Patterns in Nielsen's Latest Gauge Reports

Cable Gains Share for Second Consecutive Month in Six-Month-High Finish, Boosted...

25/06/2026

Jeopardy!, Wheel of Fortune Tap Clear-Com for Comms Upgrade

Share Copy link Facebook X Linkedin Bluesky Email...

25/06/2026

ESC 2026 - Big Blue Marble cloud-based delivery powers fl...

The Eurovision Song Contest 2026 in Vienna was a significant success for the Austrian public broadcaster ORF. In Austria, more than 1.5 million viewers tuned in...

25/06/2026

WISYCOM DELIVERS NEW LEVELS OF RF EFFICIENCY AND INFRAST...

Wisycom has further strengthened its ecosystem of professional wireless solutions with the MPR60 Wideband IEM/IFB Receiver with expanded multichannel IFB mode, ...

25/06/2026

Ease Live Powers New Interactive Experience for Rally TV...

Ease Live, the interactivity expert, today announced that its graphics overlay platform is powering a new interactive experience on Rally.TV, the official video...

25/06/2026

VFX History: the origin of After Effects

VFX History: the origin of After Effects Graham Quince June 25, 2026 0 Comments Before it was Adobe, it was CoSA. This is the VFX history of Adobe Aft...

25/06/2026

Creative Remote to Open London Offline Facility

Creative Remote, the provider of remote and hybrid offline editing infrastructure, today announced the opening of 41, its new offline edit facility located at 4...

25/06/2026

Rise Announces 2026 Worldwide Mentoring Cohorts Supportin...

Rise, the award-winning advocacy group for gender diversity in the broadcast and media technology sector, is pleased to announce the global mentoring cohort for...

25/06/2026

Emergent Partners with ROCKET to Expand Canadian Operatio...

Emergent, a pioneer in browser-based, AI-enhanced content production environments, today announced a strategic partnership with ROCKET, a premier media-centric ...

25/06/2026

Mobile Television Group Launches MTVG Full-Stack Production Platform

Share Copy link Facebook X Linkedin Bluesky Email...

25/06/2026

NAB Updates FCC on ATSC 3.0 Alerting Advances

Share Copy link Facebook X Linkedin Bluesky Email...

25/06/2026

Tegna Elevates Four Executives to Senior VP

Share Copy link Facebook X Linkedin Bluesky Email...

25/06/2026

The Ultimate Summer Sale Pairing: Steam Sale Meets GeForce NOW Discounts

Summer savings are heating up. From the Steam Summer Sale to GeForce NOW membership discounts, this week's GFN Thursday delivers double the deals and more w...

25/06/2026

June 24, 2026

Immune molecule may drive excessive drinking in alcohol use disorder Scripps Research scientists showed that blocking an immune molecule tied to inflammation r...

24/06/2026

Nielsen's Q1 2026 Ad Supported Gauge

Streaming sets record high of 46.6% of ad supported TV viewing, driven by Super Bowl and Winter Olympics; overall share of ad supported TV remains steady NEW Y...

24/06/2026

FCC Flooded with Nearly 28K Comments Regarding Its Probe of 'The View'

Share Copy link Facebook X Linkedin Bluesky Email...

24/06/2026

Hearst Television Brings Ad Addressability to Local Broadcast TV

Share Copy link Facebook X Linkedin Bluesky Email...

24/06/2026

FCC Raises $3.5 Billion in AWS-3 Wireless Auction

Share Copy link Facebook X Linkedin Bluesky Email...

24/06/2026

RE:Vision Effects Announces Twixtor Standalone v 8.1 and a Sale!

RE:Vision Effects Announces Twixtor Standalone v 8.1 and a Sale! Brie Clayton June 24, 2026 0 Comments Twixtor v8.1 Standalone adds support for variab...

24/06/2026

Dreamtek Uses Full Blackmagic Workflow for Vercel Next JS Event

Dreamtek Uses Full Blackmagic Workflow for Vercel Next JS Event Brie Clayton June 24, 2026 0 Comments Blackmagic cameras, switchers, routers, recorder...

24/06/2026

Chyron LIVE Unveils New Features: Haivision StreamHub Integration, SCTE-35 Ad Insertion, and Refined Switching Tools

Chyron LIVE Unveils New Features: Haivision StreamHub Integration, SCTE-35 Ad In...

24/06/2026

Mapping an Education

Mapping an Education How composer Chloe Clarke Smith navigated her Boston Conservatory experience and brought new meaning to her work June 24, 2026 By Sara...

24/06/2026

The Next Act

The Next Act Dean Krisha Marcano's vision for a connected Theater Division, and the fund making it possible June 24, 2026 Photo by Eric Antoniou The Or...

24/06/2026

Announcing STAGES Magazine 2026

Announcing STAGES Magazine 2026 Marking a decade since Boston Conservatory and Berklee College of Music joined forces, this issue spotlights some of the groun...

24/06/2026

Rede Legislativa Chooses Appear to Support Brazil TV Ver...

In Brazil's TV 3.0 Trials, Appear's X5 is transporting live signals from Bras lia to S o Paulo over the public internet using secure, reliable next-gene...

24/06/2026

Mediaproxy partners with HVS for US broadcast market

Melbourne, Australia - 24 June 2026: Mediaproxy, the global standard for software-based IP compliance monitoring and multiviewing solutions, has named Heartland...

24/06/2026

Gray Media Launches Political 360 Digital Advertising Solution

Share Copy link Facebook X Linkedin Bluesky Email...

24/06/2026

Walmart to Pay $1.4 Billion to Acquire Ad Tech Firm Vibe.co

Share Copy link Facebook X Linkedin Bluesky Email...

24/06/2026

FCC Flooded with Nearly 28K Comments on 'The View'

Share Copy link Facebook X Linkedin Bluesky Email...

24/06/2026

First Rush Brings SDI Multicam ProRes Recording to Apple Silicon Macs

First Rush Brings SDI Multicam ProRes Recording to Apple Silicon Macs Brie Clayton June 23, 2026 0 Comments First Rush is a native macOS application d...

24/06/2026

Vertical Drama Beneath Crimson Sails Created with Blackmagic Design

Vertical Drama Beneath Crimson Sails Created with Blackmagic Design Brie Clayton June 23, 2026 0 Comments Thunder Child Productions relies on cameras&...

24/06/2026

Seven paradoxes shaping the next era of media production - Episode 1

Why Dynamic Media Facilities Matter In this series, we explore the technologies, architectures and operational realities shaping modern media operations. Along ...

23/06/2026

Case Study: YES Networks IP Transition Expands Production Possibilities and Redefines Workflows

When we began planning our transition from an SDI-based infrastructure to a new ...

23/06/2026

Imagine Communications Appoints Greg Garmon as SVP, Americas Video Sales

Imagine Communications has announced the appointment of Greg Garmon as Senior Vice President, Americas Video Sales. Garmon will oversee account growth and busin...

23/06/2026

Snap Promotes Emma Wakely to Head of Sports and Media Partnerships, Americas

Snap has promoted Emma Wakely to Head of Sports and Media Partnerships, Americas, succeeding Anmol Malhotra, who has been elevated to Global Head of Content and...

23/06/2026

YES Network and Gotham Sports App to Air MI New York Major League Cricket Matches

YES Network and The Gotham Sports App will air MI New York's Major League Cr...

23/06/2026

HAND Issues Persistent Digital IDs to 2026 NBA Draft Class

The Universal Talent Identifier (HAND) has issued HAND IDs to 34 top projected prospects in the 2026 NBA Draft class, including AJ Dybantsa, Cameron Boozer, and...