Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

18/06/2026

Hollywood Filmmakers Launch AI Production Platform Cascade

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

ATVA Blasts Deltavision Media for Demanding 'Egregious' Retrans Fees

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

MediaKind Completes Merger with Harmonic's Video Business

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

AJA Unveils Io Xpand Thunderbolt 5 Expansion Chassis At InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

Linda May Han Oh Receives Guggenheim Fellowship

Linda May Han Oh Receives Guggenheim Fellowship The Berklee professor, bassist, and composer will use the fellowship to debut Dreams of Knowing, an interdisci...

18/06/2026

Sync and Stream: GeForce NOW Connects to Members' Game Libraries Across Devices

Play favorite titles from popular game libraries, keep progress synced and jump ...

18/06/2026

At Cannes Lions, NVIDIA Partners Reshape Advertising and Marketing With AI

The digital era gave the advertising and marketing industry speed; the AI era is giving it autonomous operations. For companies building next-generation techn...

17/06/2026

EVS Achieves EcoVadis Gold Medal, Ranking in Top 5% of Companies Globally

EVS has announced it has received the EcoVadis Gold Medal for sustainability performance, ranking among the top 5% of companies globally in the Technology/Mid-S...

17/06/2026

Chyron Releases Weather 2.4 with Updated DataFlow Module

Chyron has released Weather 2.4, an update to its weather suite for broadcasters and meteorologists. The release focuses on enhancements to the DataFlow module,...

17/06/2026

VSiN Launches Best Bets TV, a 24/7 FAST Channel

VSiN, The Sports Betting Network, has announced the launch of Best Bets TV, a free ad-supported streaming TV (FAST) channel. The 24/7 channel is currently avail...

17/06/2026

Snapchat and Team Whistle Launch World Cup Creator Program Across Miami, New York, and Los Angeles

DAZN's Team Whistle and Snap Inc. have announced a creator program centered ...

17/06/2026

LiveU Supporting Broadcast and Public Safety Operations for Summer of Soccer

LiveU is providing video transmission technology for broadcasters, production companies, and public safety agencies across North America's busy Summer of So...

17/06/2026

NATAS Announces Two New Board Members and COO Promotion

The National Academy of Television Arts and Sciences (NATAS) has announced that Laurens Grant and Jacob Ullman have joined its Board of Directors. Chief of Staf...

17/06/2026

Akta Brings AI-First Cloud Video Workflows to Oracle Cloud Infrastructure

Akta, the AI-First SaaS video platform for modern broadcast and streaming operations, today announced that its video platform is now generally available on Orac...

17/06/2026

SVG Students to Watch: Nicholas Stafford, Texas Southern University

This recent graduate from Houston found inspiration in technical directing and now eyes a future career in sports production...

17/06/2026

InfoComm 2026: Audio-Technica Brings New Wireless, Networked Audio, and Certified Conferencing Solutions

Audio-Technica (booth C7959) arrives at InfoComm 2026 in Las Vegas with a slate ...

17/06/2026

SNS Positions AI Suite as On-Premise Answer to Video Indexing and Search Demands

SNS has published a guide addressing growing demand for AI-powered video indexing, transcription, facial recognition, and searchable metadata across media libra...

17/06/2026

Providius Announces Providius Direct for Network Troubleshooting on NETGEAR Pro AV Switches

Providius has announced Providius Direct, a workflow for investigating network i...

17/06/2026

NEP Launches NEP Platform Software Orchestration System, Selects Bridge Technologies VB440

NEP Group has announced the commercial availability of NEP Platform, a software ...

17/06/2026

BBright White Paper: SRT vs. RIST for Professional Video Contribution and Distribution

A new white paper examining Secure Reliable Transport (SRT) and Reliable Interne...

17/06/2026

Bango Research: 51% of Gen Z Say Highlights Are Replacing Live Games

New research from subscription bundling platform Bango finds that younger sports fans are increasingly consuming sport through highlights, clips, and social med...

17/06/2026

SMPTE Opens Entire Standards Library to Public at No Cost

SMPTE has announced that its complete Standards catalog is now freely available to the global media technology community, including all published SMPTE Standard...

17/06/2026

MediaKind Completes Merger with Harmonic's Video Business, Creating Independent Video Infrastructure Heavyweight

Harmonic has completed the sale of its Video Business to MediaKind for $145 mill...

17/06/2026

Omaha Productions to Produce 2026 World Series of Poker Main Event Coverage for ESPN

Omaha Productions will produce the 2026 World Series of Poker (WSOP) in Las Vega...

17/06/2026

NABLF Graduates Its 2026 Broadcasting Leadership Training Class

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

ABC and ESPN Score Most-Watched NBA Finals Since 1998

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

Smartly, Roku Bring Social Performance Tools to CTV

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

Pliant Technologies Debuts Crewcom Flex at InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

The Immersive Supervisor Emerges as Hollywood's Next Production Role

The Immersive Supervisor Emerges as Hollywood's Next Production Role Brie Clayton June 17, 2026 0 Comments Above image: On set live immersive revi...

17/06/2026

Vertical Musical Playback Shot with Blackmagic PYXIS 6K

Vertical Musical Playback Shot with Blackmagic PYXIS 6K Brie Clayton June 17, 2026 0 Comments Large format sensor and DaVinci Resolve workflow used fo...

17/06/2026

DAZ 3D Launches New Game-Ready Character Assets Built for Modern Engines and Production Workflows

DAZ 3D Launches New Game-Ready Character Assets Built for Modern Engines and Pro...

17/06/2026

Spectrum Awards $1.1 Million in Digital Education Grants

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

XR Sports Alliance Adds New Members

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

AIMS Launches Free Online IPMX Training Series

Share Copy link Facebook X Linkedin Bluesky Email...

17/06/2026

Kiloview Partners with SFM to Expand AV-over-IP Solutions...

Montr al, Quebec, June 11, 2026 Kiloview, a leading provider of AV-over-IP and NDI -based video transmission solutions, today announced a distribution partner...

17/06/2026

Kiloview Launches U4 IP Video Dock Bringing Professional...

Changsha, China, June 15, 2026 Kiloview officially announced the launch of U4 IP Video Dock, a compact IP video decoder and output dock designed to bring prof...

17/06/2026

France Advances Europe's AI Future With NVIDIA Technologies

A year ago at NVIDIA GTC Paris at VivaTech, France laid out plans to advance local AI - from new AI factories and national compute capacity to open frontier mod...

17/06/2026

Half Year, Half Off! Ivory II Summer Sale is Here

Save 50% on All Ivory II Pianos and Collections Were halfway through the year, and we're cutting prices in half for the Ivory II Summer Sale! For a limite...

17/06/2026

Visibility builds credibility - the tools you use every day...

Visibility builds credibility - the tools you use every day, now visible on your LinkedIn profile Published on Jun 17, 2026 Categories: Company News, Product ...

17/06/2026

June 16, 2026

Calibr-Skaggs awarded $5.1M by NIH to develop long-acting hepatitis B virus therapy A new program aims to replace a daily HBV drug with once-monthly or even qua...

16/06/2026

Neumann MT 48 Receives Major Firmware 2.0 Update

Neumann.Berlin has released firmware version 2.0 for the MT 48 audio interface, adding plugin compatibility, expanded Dante networking options, broadcast encode...

16/06/2026

TVNewsCheck Opens Nominations for 2027 Women in Technology Awards

TVNewsCheck has announced that nominations are now open for its 2027 Women in Technology Awards, to be presented at NAB Show 2027 on Tuesday, April 6 in the Med...

16/06/2026

Clear-Com Introduces Avalon IP Intercom Platform

Clear-Com has announced Avalon, a 1RU IP intercom platform for broadcast, live events, and production environments. Designed for IP-only workflows, Avalon suppo...

16/06/2026

SNS EVO Enables Remote and Distributed Video Editing Workflows

SNS has published a guide to remote video editing workflows using its EVO shared storage platform and companion tools, covering use cases ranging from home edit...

16/06/2026

Richmond Flying Squirrels Deploy Grass Valley LDX 110 Cameras at CarMax Park

Grass Valley has announced that the Richmond Flying Squirrels, a Minor League Baseball affiliate of the San Francisco Giants, have deployed five Grass Valley LD...

16/06/2026

AIMS Launches Free Official IPMX Training Series Online

The Alliance for IP Media Solutions (AIMS) has announced the launch of the Official IPMX Training Series, a free online program covering the design, configurati...

16/06/2026

Swerve Womens Sports Announces Distribution Deals with Fubo, Plex, Amazon Fire TV, and Anoki AI

Swerve TV has announced distribution agreements with Fubo, Plex, Amazon Fire TV,...

16/06/2026

ATP and TikTok Expand Global Content Partnership

ATP and TikTok have announced an expansion of their global content partnership, extending the ATP's TikTok hub powered by TikTok GamePlan to cover all nine ...

16/06/2026

FOX Sports Turns Los Angeles Pico Lot Into Its FIFA World Cup Production Nerve Center

Network's LA facility serves as the heart of a sprawling operation built to ...

16/06/2026

300+ Records a Day, 150 TB Daily, and a Relentless Content Avalanche: Inside FOX Sports' World Cup Media Engine

At Pico, the network's media-management team is supporting a flood of HBS fe...