Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

28/05/2026

Cool Concentric Text for Cavalry

Cool Concentric Text for Cavalry Simon Ubsdell May 27, 2026 0 Comments In this tutorial I'll show you how to make this popular text effect that is...

27/05/2026

Telestream Appoints Benjamin Desbois as CEO, Effective July 1

Telestream has announced that its Board of Directors has appointed Benjamin Desbois as Chief Executive Officer, effective July 1, 2026. Desbois, currently Teles...

27/05/2026

FOX MLB Leads Live-Event Categories; ESPN Is Tops Overall at 47th Annual Sports Emmy Awards

ESPN garnered 10 awards; NBC's Sunday Night Football received the Outstandin...

27/05/2026

Matrox Video Marks 50th Anniversary, Announces New Product Launch for June

Matrox Video is celebrating its 50th anniversary, marking five decades of operations from its headquarters in Montreal, Canada. Founded in 1976, the company has...

27/05/2026

MLB Announces Fan Engagement Initiatives for Americas 250th Anniversary

Major League Baseball has announced a series of initiatives tied to America's Semiquincentennial, including a national marketing campaign, Fourth of July br...

27/05/2026

Advanced Systems Group Hires Brian Gross as Account Manager for Audio Team

Advanced Systems Group (ASG) has announced that Brian Gross has joined the company as an Account Manager on its Audio team, based in the Burbank office. He will...

27/05/2026

Nielsen Research: Hispanic Fans, Asian Markets Drive Global Soccer Audience Ahead of World Cup 2026

Nielsen has released new research on soccer fandom ahead of the FIFA World Cup 2...

27/05/2026

ESL FACEIT Group Debuts First Ever Esports Vertical Stream Co-Developed With TikTok

ESL FACEIT Group (EFG) has unveiled a new partnership with TikTok to bring broad...

27/05/2026

Two Weeks Away: FIFA Outlines Production Plans for Highly Anticipated North American-Based World Cup

FIFA's Oscar Sanchez gives a deeper look to how this tournament will be cove...

27/05/2026

SVG Students To Watch: Maggie Lynn, Virginia Tech

The soon-to-be senior from Charlottesville is building her skills in replay, TD, and even creative content for HokieVision and its ACC Network productions In t...

27/05/2026

A Global Festival of Football: FOX Sports Illustrates Strategy to Bring Every FIFA Mens World Cup Match to the U.S. Audience

FOX Sports' Mike Davies breaks down the vision for this summer's showcas...

27/05/2026

Top-Tier Storytelling: Host Broadcast Services Works at Capturing the Atmosphere of the FIFA Mens World Cup

HBS's Paul King, FIFA's Oscar Sanchez preview how the masses at home wil...

27/05/2026

Matt Gangl & Pete Macheska on FOX MLBs Huge Night and an Unforgettable Postseason Run

FOX's MLB coverage dominated the night at the 47th Annual Sports Emmy Awards...

27/05/2026

FOXs Mike Davies and Team on Outstanding Technical Team Win for 2025 World Series

One of the most memorable Postseasons in baseball history would have had no memo...

27/05/2026

NBC Sports Rob Hyland Reflects on an Unforgettable Sunday Night Football Season

NBC's Sunday Night Football is among the most decorated and most watched programs in the history of television. It added to its jam-packed trophy case on Tu...

27/05/2026

Prime Videos John Ward and Mike Francis on Groundbreaking NBA on Prime Video Studio

The 2026 Sports Emmys marked a watershed moment for Prime Video Sports. After bu...

27/05/2026

Countdown to FIFA World Cup 2026: SVG Launches SportsTechLive Blog in Lead-up to Winter Games

With the Opening Match just over two weeks away, the entire sports-production-te...

27/05/2026

L3Harris Introduces the XL Converge 300P Portable Public Safety Radio

The XL Converge 300P radio system emerges with a groundbreaking feature set enhancing the mission-critical communications of public safety, federal and critica...

27/05/2026

Modernizing Public Safety Communications

Pairing Two47 MCX software with existing LTE networks means tailored system upgrades that can save time, money and lives....

27/05/2026

L3Harris Strengthens Global Solid Rocket Motor Supply Chain With New PAC-3 Propulsion Supplier

PAC-3 MSE offers improved range, speed, and maneuverability, making it an effect...

27/05/2026

Brightcove Adds New Features to Its AI Suite for Video Advertising

Share Copy link Facebook X Linkedin Bluesky Email...

27/05/2026

Star Trek VFX: Recreating John Knoll's Iconic Warp Stars without a Slitscan Camera

Star Trek VFX: Recreating John Knoll's Iconic Warp Stars without a Slitscan ...

27/05/2026

Adventure World Uses Blackmagic Replay for Marine Live

Adventure World Uses Blackmagic Replay for Marine Live Brie Clayton May 27, 2026 0 Comments Large screen displays and slow motion replays dynamically ...

27/05/2026

Berklee Alumna and Assistant Professor Olivia Prez-Collellmir to Premiere Original Work at Gaud Centennial in Barcelona

Berklee Alumna and Assistant Professor Olivia P rez-Collellmir to Premiere Origi...

27/05/2026

Gravity Media Expands Into Creative Services With New Agency

Share Copy link Facebook X Linkedin Bluesky Email...

27/05/2026

Tegna Names Patrick Paolini as CEO

Share Copy link Facebook X Linkedin Bluesky Email...

27/05/2026

Telestream Taps Company Vet Benjamin Desbois as CEO

Share Copy link Facebook X Linkedin Bluesky Email...

27/05/2026

HDR10+ Technologies to Launch Eclipsa Video Certification Program

Share Copy link Facebook X Linkedin Bluesky Email...

27/05/2026

ATSC to Gather in Washington Next Week for Annual Meeting

Share Copy link Facebook X Linkedin Bluesky Email...

27/05/2026

Telestream Appoints Benjamin Desbois as Chief Executive O...

Co-founder Dan Castles to transition to Executive Chair; internal promotion reinforces continuity and long-term growth Telestream, a global leader in media wor...

27/05/2026

Big Blue Marble Announces First End-to-End 5G Broadcast S...

Big Blue Marble today announced that its Nakolos platform is the first end-to-end 5G Broadcast solution worldwide to implement the complete feature set introduc...

27/05/2026

Lightware Continues Its ESG Commitment Through Girls Day...

Lightware recently hosted the Girls' Day event in April at its headquarters in Budapest, welcoming students for an interactive introduction to engineering a...

26/05/2026

Matrox Video Marks 50 Year Milestone

Share Copy link Facebook X Linkedin Bluesky Email...

26/05/2026

Roku Expands Premium Subscriptions With Fox One

Share Copy link Facebook X Linkedin Bluesky Email...

26/05/2026

Brian Gross Joins ASG's Audio Team as Account Manager

Share Copy link Facebook X Linkedin Bluesky Email...

26/05/2026

MPA Urges FCC Not to Reclassify vMVPDs

Share Copy link Facebook X Linkedin Bluesky Email...

26/05/2026

Cobalt Digital to Showcase End-to-End IPMX Ecosystem at I...

Cobalt Digital to Showcase End-to-End IPMX Ecosystem at InfoComm 2026, Making ST 2110 Easy for Pro AV blueCORE standalone processors headline solutions designe...

26/05/2026

Matrox Video Celebrates 50 Years of Innovation and Looks...

Matrox Video today celebrates its 50th anniversary, marking five decades of innovation, engineering excellence, and customer-focused evolution from its headquar...

26/05/2026

Cuez Helps ITV Studios Run Three Live Shows as One includ...

ITV Studios, the production arm of the UKs largest commercial broadcaster, has deployed the Cuez live production platform to unify the management of three back-...

26/05/2026

CETA Software releases Morpheus, an AI tool for real-time post-production project oversight

CETA Software releases Morpheus, an AI tool for real-time post-production projec...

26/05/2026

Ikegami Accelerates Motorsport Broadcast Innovation with...

In a move set to redefine motorsports coverage across the Asia Pacific region, Ikegami Electronics announces that Two Wheels Motor Racing Sdn Bhd (TWMR), a lead...

26/05/2026

Netflix Releases Official Trailer and Poster for 'The Root Of The Game,' A Documentary Series Premiering June 20

Back to All News Netflix Releases Official Trailer and Poster for The Root Of T...

26/05/2026

'MED,' Production Begins on Netflix's First Medical Drama From Brazil, Starring Clara Moneke

Back to All News MED, Production Begins on Netflixs First Medical Drama From Br...

26/05/2026

Netflix Announces Five New Brazilian Productions and Expands Its Presence at Rio2C 2026

Back to All News Netflix Announces Five New Brazilian Productions and Expands I...

26/05/2026

Netflix Celebrates the Release of the New Animated Series 'Due Spicci' With a Major Event at Circo Massimo in Rome

Back to All News Netflix Celebrates the Release of the New Animated Series Due ...

26/05/2026

Made In New Mexico: Building The Boroughs' From the Ground Up

Back to All News Made In New Mexico: Building The Boroughs' From the Ground Up A photo from The Boroughs.' (Courtesy of Netflix 2026) Entertainme...

26/05/2026

Broadcast Pix Introduces ONix Pro Control Panel

Smart Production Control. Total Confidence. Tyngsborough, MA, May 27, 2026 - Broadcast Pix today announced the ONix Pro Control Panel, its most advanced hard...

26/05/2026

NVIDIA Vera CPU Is Packing a Heavy-Hitting Punch' Against Competition

The shift to agentic AI creates a new CPU requirement for the AI factory: fast cores, massive memory bandwidth and the ability to sustain high performance when ...

25/05/2026

SVG All-Stars: LJ Helbig, Senior Manager, Broadcast Engineering, FloSports

The former University of Wyoming wrestler is essential in helping the rapidly growing streamer delivers more than 50,000 live events per year The sports-produc...