Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

13/02/2026

Rai Selects Imagine Selenio Network Processor for IP Migration

Share Copy link Facebook X Linkedin Bluesky Email...

13/02/2026

CIMM Details Research Plans for 2026 and New Board Appointments

Share Copy link Facebook X Linkedin Bluesky Email...

13/02/2026

Teradek Unveils RF-A Auto Switcher

Share Copy link Facebook X Linkedin Bluesky Email...

13/02/2026

Spectrum Launches 'Invincible Wifi'

Share Copy link Facebook X Linkedin Bluesky Email...

13/02/2026

Actus Digital to Introduce Actus X Platform Enhancements At NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

13/02/2026

Sennheiser Wireless Spectera Solution Tackles Super Bowl LX With Ease

Share Copy link Facebook X Linkedin Bluesky Email...

13/02/2026

Nate Bargatze to Receive 2026 NAB Television Chairman's Award

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Chyron Merges Live Web Content and CG Graphics with PRIME 5.3

Chyron unveils PRIME 5.3, the latest software release of the company's powerful engine for live production graphics. PRIME 5.3 delivers the first official i...

12/02/2026

SVG New Sponsor Spotlight: Interra Systems' Anupama Anantharaman on Protecting Live Sports Quality Across IP and OTT Workflows

The vendor's VP of Product Management explains how quality assurance, monito...

12/02/2026

LTN Makes Key Appointments and Introduces New Technology Organization

LTN announces the appointment of three experienced executives to lead its new Technology organization: Michal Miskin-Amir as EVP and Head of Technology, Jonatha...

12/02/2026

Riedel Opens Kuala Lumpur Office to Strengthen Global 24/7 Software and IT Support

Riedel Communications has officially opened a new office in Kuala Lumpur, Malays...

12/02/2026

NATO Upgrades Brussels HQ Broadcast Studio with Grass Valley LDX 135 Cameras

Grass Valley has won a competitive NATO-wide tender to provide the new camera system for NATO's main broadcast studio at its Brussels headquarters. The proj...

12/02/2026

Canon Announces Big Game Broadcast Lens Use Data

Canon U.S.A announces that the vast majority of broadcast lenses utilized on the NBC live broadcast for the Big Game between New England and Seattle on Sunday w...

12/02/2026

Three-Time Grammy-Winner Ludacris to Headline Performances at NBA All-Star 2026 in LA

The National Basketball Association (NBA) and NBC Sports announce the entertainm...

12/02/2026

IOC Awards Broadcast Rights in Middle East and North Africa to beIN MEDIA GROUP

The International Olympic Committee (IOC) announces that beIN MEDIA GROUP ( beIN ), the leading global sports, entertainment and media organisation, has secured...

12/02/2026

Big 12 Conference Unveils ASB GlassFloor for Upcoming Tournaments in March

The Big 12 Conference and ASB GlassFloor introduces a full LED video sports floor that will debut at the 2026 Phillips 66 Big 12 Men's and Women's Baske...

12/02/2026

TNDV Showcases Aspiration35 at National Religious Broadcasters Convention 2026

Continuing its commitment to serving the faith-based broadcast and live event community, mobile production company TNDV, a division of Live Media Group, will hi...

12/02/2026

Marshall Electronics POVs Power Hidden-Camera Investigations on German TV Series, ACHTUNG ABZOCKE

The production team of the long-running German investigative series Achtung Abz...

12/02/2026

Vizrt Launches Sports Production Bundles to Empower US Students to Produce Like the Pros

Vizrt announces the launch of four Campus Stadium Production Bundles, designed t...

12/02/2026

LiveU Spotlights Three Broadcast Priorities with Digital-First, Workflow Automation, IP Contribution Resilience

At NAB Show, LiveU will showcase its broadest IP-video EcoSystem to date, design...

12/02/2026

Follow the Money, Episode 5: Analyzing Sports Media Deals With Sam McCleery and PwC's Lori Bistis

Welcome to the Sports Video Group's new interview series, Follow the Money, ...

12/02/2026

NBC Sports Engineers a Super Bowl Transmission Plan That Reached From Alcatraz to Stamford - and Out Onto San Francisco Bay

400 Gbps of bandwidth, layered redundancy, and mobile-first connectivity powered...

12/02/2026

L3Harris' VAMPIRE System Successfully Fires Thales' Belgian-Made 70 mm Rockets

L3Harris' VAMPIRE system fires Thales Belgian-made 70 MM rocket from an FZ60...

12/02/2026

LTN Names 3 Executives to Lead Technology Group

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

On the Ice, There's a Third Team at Work

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Mo Rocca to Receive the 2026 LABF Insight Award at NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Marshall Electronics POVs Power Hidden Camera Investigat...

The production team of the long-running German investigative series Achtung Abzocke recently upgraded its cameras for the show's 12th season. The objectiv...

12/02/2026

Bitmovin Appoints Ian Baglow as Co-Chief Executive Office...

Leading provider of video streaming solutions, Bitmovin, has appointed Ian Baglow as Co-CEO alongside existing CEO and Co-Founder Stefan Lederer. Under this str...

12/02/2026

Vizrt Launches Sports Production Bundles to Empower US St...

Vizrt, a leading viewer engagement platform and a trusted expert in live production technologies, today announces the launch of four Campus Stadium Production B...

12/02/2026

Ailanto and Cubbit launch sovereign cloud storage for Swi...

Strategic agreement to deliver S3 cloud storage in Switzerland with full data sovereignty and local control including at the level of individual cantons plu...

12/02/2026

Mad About Video counts on Lightware MX2 matrix switcher a...

Mad About Video is a leading specialist in video for live events and installations throughout Malta. In operation since 2011, it has evolved from a company focu...

12/02/2026

JAGGAER supports Betsson Group in further strengthening i...

JAGGAER, a global leader in digital procurement and supplier collaboration solutions, today announced the successful delivery of a procurement digitalization pr...

12/02/2026

LiveU Spotlights Three Broadcast Priorities at NAB Show 2...

At NAB Show, LiveU will showcase its broadest IP-video EcoSystem to date, designed to help broadcasters and content creators embrace digital first operations, d...

12/02/2026

Spectrum News Acquires New England Cable News

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Vizrt Unveils Campus Stadium Production Bundles

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Hulu + Live TV Adds Fubo Sports Network to Channel Line-up

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

FCC To Hold Open Commission Meeting on Feb. 18

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Ralph M. Oakley to Receive NAB's Chuck Sherman TV Leadership Award

Share Copy link Facebook X Linkedin Bluesky Email...

12/02/2026

Netflix unveils the trailer for 'That Night'

Back to All News Netflix unveils the trailer for That Night Entertainment 12 February 2026 GlobalSpain Link copied to clipboard WATCH THE TRAILER DOWNLOA...

12/02/2026

NVIDIA DGX Spark Powers Big Projects in Higher Education

At leading institutions across the globe, the NVIDIA DGX Spark desktop supercomputer is bringing data center class AI to lab benches, faculty offices and studen...

12/02/2026

Leading Inference Providers Cut AI Costs by up to 10x With Open Source Models on NVIDIA Blackwell

A diagnostic insight in healthcare. A character's dialogue in an interactive...

12/02/2026

GeForce NOW Turns Screens Into a Gaming Machine

The GeForce NOW sixth-anniversary festivities roll on this February, continuing a monthlong celebration of NVIDIA's cloud gaming service. This week brings ...

12/02/2026

February 11, 2026

TIME100 Health list features Scripps Research Professor Darrell Irvine Irvine is recognized for his work in empowering the immune system to fight disease, which...

11/02/2026

FYI: Phone Support Maintenance

FYI: Phone Support Maintenance One thing we pride ourselves on here at Utah Scientific is our 24-hour support included with our signature 10-year hardware warra...

11/02/2026

Bitmovin Appoints Ian Baglow as Co-Chief Executive Officer

Leading provider of video streaming solutions, Bitmovin, has appointed Ian Baglow as Co-CEO alongside existing CEO and Co-Founder Stefan Lederer. Under this str...

11/02/2026

Paramount and CBS Partner to Air UFC 326

Paramount and the CBS Television Network will partner to air UFC 326: HOLLOWAY vs. OLIVEIRA 2 live on Saturday, March 7, from T-Mobile Arena in Las Vegas, mar...

11/02/2026

MLB.TV Launches on ESPN Beginning February 10

Beginning February 10, fans can buy MLB.TV on ESPN, a new milestone in one of sports media's longest-standing partnerships. ESPN becomes the new streaming h...

11/02/2026

Fubo Sports Network Launches on Hulu + Live TV

Fubo Sports Network is available to Hulu's Live TV subscribers in the core $89.99 a month subscription plan, which also includes full access to the entire H...

11/02/2026

Rai Selects Imagine Selenio Network Processor for IP Migration

Following a competitive public tender process, Rai (Radiotelevisione Italiana), the national public broadcasting company of Italy, has awarded Imagine Communica...