Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

12/05/2026

Beyond the Hype: A Strategic Post-Hoc Analysis of NAB 2026

Beyond the Hype: A Strategic Post-Hoc Analysis of NAB 2026 If NAB Show 2026 had an underlying theme, it was a quiet, industry-wide pivot from the high-energy sp...

12/05/2026

G&D and CT Square Establish Joint Venture for India and Middle East Distribution

Guntermann and Drunck (G&D), a Panoptec Technologies Group company, and CT Square, led by Chandresh Shah, have announced a joint venture to distribute G&D and V...

12/05/2026

Telemundo Reports Record Advertiser Demand for FIFA World Cup 2026 Coverage

With 30 days until the start of the FIFA World Cup 2026, Telemundo, the exclusive Spanish-language home of the tournament in the United States, has announced th...

12/05/2026

NHLs Stanley Pup Returns for Third Year on June 8

The NHL has announced the return of Stanley Pup for its third consecutive year, a 90-minute special featuring adoptable rescue dogs competing on a miniature rin...

12/05/2026

NBCUniversal Outlines Ad Technology and Content Strategy at 2026 Upfront

NBCUniversal presented its 2026 Upfront to advertisers at Radio City Music Hall, detailing upcoming programming across NBC, Peacock, Bravo, and Versant properti...

12/05/2026

FOX Sports Funds Fandom Research Initiative at Harvard Kennedy School

FOX Sports has announced funding for the Fandom and Social Connection Initiative at Harvard Kennedy School's Shorenstein Center on Media, Politics, and Publ...

12/05/2026

TNDV and Live Media Support NCAA Final Four Weekend for 13th Consecutive Year

TNDV and Live Media, both divisions of Live Media Group, supported live broadcast coverage around NCAA Final Four weekend in Indianapolis, including the March M...

12/05/2026

European Football Alliance Signs Distribution Agreement with Fubo Sports Network

The European Football Alliance (EFA) has announced a content distribution agreement with Fubo Sports Network, the free ad-supported streaming TV (FAST) channel ...

12/05/2026

ESPN, TelevisaUnivision Ink Deal to Produce Additional Spanish-Language Super Bowl LXI Telecast

For the first time, Spanish-speaking fans in the U.S. will have two separate tel...

12/05/2026

CP Communications Manages RF Spectrum for Kentucky Derby Week

CP Communications led a comprehensive spectrum management initiative on behalf of Churchill Downs during Kentucky Derby week, coordinating RF assets across the ...

12/05/2026

LiveU and DRONERESPONDERS Announce Partnership for Public Safety Drone Video Transmission

LiveU has announced a strategic partnership with DRONERESPONDERS, a 501(c)3 non-...

12/05/2026

BMC TV Selects Open Broadcast Systems 5G Flyaway for Horse Racing Transmissions

Open Broadcast Systems has announced that BMC TV, a specialist in IP transport of broadcast content, has selected the Open Broadcast Systems 5G Flyaway solution...

12/05/2026

NEP Europe to Deliver Broadcast Solutions for Eurovision Song Contest 2026 in Vienna

NEP Europe, part of NEP Group, has announced it will deliver broadcast solutions...

12/05/2026

Grass Valley Expands Collaboration with Ravensbourne University London on Media Production Education

Grass Valley has announced continued collaboration with Ravensbourne University ...

12/05/2026

Stats Perform Launches Opta Pulse AI-Assisted Video Creation Platform

Stats Perform has announced the launch of Opta Pulse, an AI-assisted video creation and distribution platform for leagues, rights holders, and broadcasters. The...

12/05/2026

FOX Sports and Sesame Workshop Announce Collaboration for FIFA World Cup 2026

FOX Sports has announced a collaboration with Sesame Workshop to integrate Sesame Street characters into FOX Sports' FIFA World Cup 2026 programming. Conten...

12/05/2026

Little Show, Giant Heart: NHL in ASL Earns Two Sports Emmy Nominations in 2026

To date, NHL Productions has produced 19 broadcasts with commentary in American Sign Language NHL in ASL (American Sign Language) may be just one show, but the...

12/05/2026

At Its Annual Brandcast, YouTube Aims To Show Why Sports Is No Longer Just the Game

Google's Brian Albert: creators, athletes, highlights, nostalgia, second-scr...

12/05/2026

Your Next Double Feature: Old Joy and Past Lives

A still from Past Lives by Celine Song, an official selection of the Premieres program at the 2023 Sundance Film Festival. (Courtesy of Sundance Institute | p...

12/05/2026

CNN Pushes Into Lifestyle Products with CNN Weather' App Launch

Share Copy link Facebook X Linkedin Bluesky Email...

12/05/2026

Marshall Electronics Cameras Elevate Elite Equestrian Com...

Tyrell Corporation, specialists in high-end live sports and entertainment broadcasts, was tasked with delivering compelling broadcast coverage of premier equest...

12/05/2026

IBC2026 registration opens as international media industr...

Registration is now open for IBC2026 as the global media, entertainment and technology community prepares to converge on the RAI Amsterdam from 11 14 September ...

12/05/2026

Ross Video Showcases Unified Production Workflows at Broa...

Ross Video, a global leader in live video production technology, will present its latest innovations and integrated production workflows at BroadcastAsia 2026, ...

12/05/2026

With over 50bn euros in assets under management and recor...

500 selected leaders from around the world across start-ups, corporates, and venture capital. Over 50bn in assets under management among attending investors, a...

12/05/2026

Randy Koeman Joins CVP as Part of Continued European Expa...

CVP, one of Europe's leading suppliers of professional video and broadcast solutions, has appointed Randy Koeman as Sales Manager, Netherlands, marking anot...

12/05/2026

Taiwan Television selects PlayBox Neo Suite multi-channel...

Taiwan Television has expanded its HD content ingest workflow capabilities by investing in the PlayBox Neo Suite - a leading multi-channel multi-server UHD/HD/S...

12/05/2026

Jigsaw24 Appoints Tom Laker and Ashish Chauhan to Support...

Jigsaw24 is strengthening its media team with the appointment of Tom Laker as Client Director and Ashish Chauhan as Project Manager. These strategic hires bring...

12/05/2026

Pebble brings future-ready playout innovation to Broadcas...

Broadcast playout leader showcases hybrid IP infrastructure and workflow automation with Videoland deployment as regional benchmark...

12/05/2026

dzjinius to Showcase All-in-one Production Management Pla...

dzjinius, the all-in-one production management platform for the media and entertainment industry, will showcase its centralised approach to production workflows...

12/05/2026

G&D and CT Square Launch New Joint Venture in India

Share Copy link Facebook X Linkedin Bluesky Email...

12/05/2026

Congress Urged to Protect Live Sports on Broadcast TV

Share Copy link Facebook X Linkedin Bluesky Email...

12/05/2026

Roku to Stream Inaugural Enhanced Games in North America

Share Copy link Facebook X Linkedin Bluesky Email...

12/05/2026

Austrian Broadcaster Selects Lawo, SLG Broadcast for Studio Upgrade

Share Copy link Facebook X Linkedin Bluesky Email...

12/05/2026

NUGEN Audio Debuts Enhanced DialogCheck at MPTS 2026

NUGEN Audio will debut Version 1.1 of DialogCheck, its intelligent dialog intelligibility and compliance tool, at the 2026 Media Production & Technology Show (S...

12/05/2026

Disguise Powers Interactive LED Castle for Laura Pausini

The reactive visuals can adapt to any stage or screen setup and comprise 50 layers of Notch content Today, Disguise announced that its high performance GX 3 m...

12/05/2026

DoPchoice Intros Light-Shaping Accessories for ARRI Omnib...

Snapgrid , Snapbag & Airglow Expand Creative Control for the ARRI Modular LED BarFollowing the launch of the ARRI Omnibar, DoPchoice introduces a dedicated li...

12/05/2026

Tribeca Festival Unveils 2026 Games Program

May 12th, 2026 Press Materials Available Here TRIBECA FESTIVAL UNVEILS 2026 GAMES PROGRAM Tribeca Celebrates Its 25th Anniversary With World Premieres and Pl...

12/05/2026

Madonna Brings The World Premiere Of Confessions II To 2026 Tribeca Festival

May 12th, 2026 Press Materials Available Here MADONNA BRINGS THE WORLD PREMIERE OF CONFESSIONS II TO 2026 TRIBECA FESTIVAL The Global Icon Will Unveil a Cine...

12/05/2026

Macarena Garca and Carlos Cuevas star in 'El Acercamiento de la Mujer Cactus y el Hombre Globo'

Back to All News Macarena Garc a and Carlos Cuevas star in El Acercamiento de l...

12/05/2026

Netflix Unveils Documentary on Norway's 26-Year Journey to This Year's World Cup

Back to All News Netflix Unveils Documentary on Norway's 26-Year Journey to...

12/05/2026

The Netflix Effect

Back to All News The Netflix Effect Ted Sarandos co-CEO Social Impact 12 May 2026 Global Link copied to clipboard Download all assets Ten years ago, Ne...

12/05/2026

News Test

This is my news content...

12/05/2026

NVIDIA and SAP Bring Trust to Specialized Agents

From finance and procurement to supply chain and manufacturing, specialized AI agents are moving into the enterprise systems where business decisions are made, ...

11/05/2026

Ross Production Services 48-Foot HyperMax-1 Truck Quickly in Demand

New production unit features not only wealth of Ross gear but also Sony cameras, Canon lenses, and a Calrec Argo S audio board...

11/05/2026

Solid State Logic Launches TCA Tour Portable Audio Production System

Solid State Logic (SSL) has announced TCA Tour, a portable fly-away audio production system built from System T components for broadcast, touring, and live prod...

11/05/2026

Riedel Communications Named Official Connectivity Integration Provider for Glasgow 2026 Commonwealth Games

Riedel Communications has announced it will serve as Official Connectivity Integ...

11/05/2026

G&D and NETGEAR AV Launch KVM-over-IP Plugin for Automated Network Configuration

Guntermann and Drunck (G&D) and NETGEAR AV have announced a plugin that automates network configuration for G&D KVM-over-IP deployments on NETGEAR AV Line infra...

11/05/2026

The Famous Group and Elite Edge Announce Unified Ownership

The Famous Group (TFG) and Elite Edge, a production studio specializing in live-action production, show opens, set design, and fabrication for sports and live e...

11/05/2026

SVG All-Stars: Odair Auger, Tech Project Manager, NBCUniversal Telemundo Enterprises

The native of S o Paulo, Brazil, will play a crucial role in the broadcaster'...

11/05/2026

Cobalt Digital to Showcase IPMX/ST 2110 Ecosystem at BroadcastAsia 2026

Cobalt Digital will exhibit its IPMX/ST 2110 product lineup at BroadcastAsia 2026 (Stand 5D2-4), anchored by the COBALT blueCORE family of standalone signal pro...