Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

24/11/2025

Sinclair Promotes Sean LaRose to VP and General Manager in Rochester

ROCHESTER, N.Y. Sinclair said it has elevated Sean LaRose, director of sales at WUHF and partner station WHAM here, to vice president and general manager, effec...

24/11/2025

AI On: 3 Ways Specialized AI Agents Are Reshaping Businesses

Editor's note: This post is part of the AI On blog series, which explores the latest techniques and real-world applications of agentic AI, chatbots and copi...

22/11/2025

Deadline Extended for 2025 Best in Market Awards

The deadline for entries for the 2025 Best in Market Awards has been extended to 23:59 PST on November 28, 2025....

22/11/2025

Clear-Com Unveils 4-Channel HelixNet Beltpack - Expanding...

Clear-Com announced the upcoming launch of its 4-Channel HelixNet beltpack, a next-generation advancement of its widely used 2-channel model. The new beltpack...

22/11/2025

Marshall Electronics Features its CV625 and CV612 PTZ Cam...

Marshall Electronics is showcasing the latest additions to its CV600 Series of PTZ cameras, the CV625 and CV612, which both feature AI track and follow capabili...

22/11/2025

QuickLink StudioPro Powers LiveConnect Live Production at...

At this year's European Respiratory Society (ERS) Congress, held at the RAI Amsterdam, LiveConnect delivered an ambitious and technically complex live produ...

22/11/2025

Professional Wireless Systems PWS Delivers Flawless RF Co...

Professional Wireless Systems (PWS), a leader in wireless frequency coordination and RF system design, provided a comprehensive wireless gear package and onsite...

22/11/2025

Telestream Introduces ARGUS v23 Featuring Live Look for R...

Telestream, a global leader in media workflow technologies, today announced the release of ARGUS v2.3, which introduces Live Look, a powerful new feature that e...

22/11/2025

Peer Software Expands Data Orchestration and Analytics Pl...

Peer Software today announced significant advancements across its enterprise data orchestration and analytics platform with new releases of Peer Global File Ser...

22/11/2025

Atomos Expands Ninja TX Series Capabilities with Powerful...

At InterBEE 2025, Atomos announces a major firmware update that brings integrated camera control to the Ninja TX GO and Ninja TX its new CFexpress-based monit...

22/11/2025

AWS Announces Elemental MediaConnect Router

Today, AWS announces the general availability of AWS Elemental MediaConnect Router, a new capability that enables broadcasters and content providers to dynamica...

22/11/2025

Rise Announces 2025 Award Winners

Rise, the award-winning advocacy group for gender diversity in the broadcast media technology sector, is delighted to announce the winners for this year's R...

22/11/2025

Lightware introduces new HC60 product line to strengthen...

Lightware, industry leader in signal management, is strengthening its Taurus UCX product family with the introduction of the new HC60 lineup. The new product li...

22/11/2025

IDX Unveils CUE-J Series Batteries

CARSON, Calif. IDX has introduced the IDX CUE-J Series battery/charger kits, including the CUE-J98, CUE-J150 and CUE-J198....

22/11/2025

NBA: More Than 60 Million Watched Games in First Month, Best in 15 Years

The NBA has released encouraging viewing and social media data that the beginnings of its $76 billion deal with NBC/Peacock, Prime Video and ESPN are paying off...

22/11/2025

FCC Sets Deadlines for Comments on New NextGen TV Proposals

WASHINGTON The Federal Communications Commission has set deadlines for comments on its newest proposals for NextGen TV, aka ATSC 3.0, with comments due on Jan. ...

22/11/2025

Seeking Advice for a New Opera, Laura Kaminsky Consulted the Experts: Her Students

Seeking Advice for a New Opera, Laura Kaminsky Consulted the Experts: Her Studen...

21/11/2025

Platinum White Paper: Appear Shares Why Media Exchange Is the Missing Link in Software-Defined Live Production

Platinum White Paper: Appear Shares Why Media Exchange Is the Missing Link in So...

21/11/2025

NWSL Championship 2025: CBS Sports To Deploy Two-Point FlyCam for Match Coverage at PayPal Park

NWSL Championship 2025: CBS Sports To Deploy Two-Point FlyCam for Match Coverage...

21/11/2025

NWSL Caps 2025 Season With Awards Show, Skills Challenge Productions

NWSL Caps 2025 Season With Awards Show, Skills Challenge ProductionsA team of 70 is on the ground in California to produce both eventsBy Mark J Burns, SVG Contr...

21/11/2025

USL and NEP Ready for Largest USL Championship Final Production Ever

USL and NEP Ready for Largest USL Championship Final Production EverThe broadcast from Tulsa, OK, will air CBS and TUDN on Saturday at 12 p.m. ETBy Jason Dachma...

21/11/2025

With Two New Teams, PWHL Boosts Production Workforce and Central Review for Season 3

With Two New Teams, PWHL Boosts Production Workforce and Central Review for Seas...

21/11/2025

Dinner and a Movie: Jared Lank on Powwow Highway and Luskinikn

Jared Lank and his mother in the '90s...

21/11/2025

L3Harris Recognizes Employee Achievements with LHX Excellence Awards

MELBOURNE, Fla., Nov. 21, 2025 - L3Harris Technologies (NYSE: LHX) has announced this year's LHX Excellence Awards, the company's most prestigious recog...

21/11/2025

FCC Proposes Upper C-Band Rules for 2027 Auction

WASHINGTON The Federal Communications Commission by a 3-0 vote opened a notice of proposed rulemaking (NPRM) to advance Congress's mandate to clear a minimu...

21/11/2025

FCC Votes to Clear at Least 100MHz of Upper C-Band Spectrum

WASHINGTON The Federal Communications Commission by a 3-0 vote adopted a Notice of Proposed Rulemaking (NPRM) to advance Congress's mandate to clear a minim...

21/11/2025

Spectrum Expands 4K Content to Apple TV 4K and Roku Devices

STAMFORD, Conn. Charter Communications' Spectrum brand has expanded the range of devices that can offer 4K content on the Spectrum TV app to compatible Appl...

21/11/2025

Study: Salaries in Content, Connectivity Industries Continue To Grow

NAPERVILLE, Ill. Media industry employers are continuing their multiyear trend of increasing salaries for all worker segments but lag general industry raises, s...

21/11/2025

NAB Opens Nominations for 2026 Technology Awards

WASHINGTON The National Association of Broadcasters said it is accepting nominations for the 2026 NAB Technology Awards, honors that recognize excellence in bro...

21/11/2025

AAT Introduces Automated RF Line Analysis

American Amplifier Technologies has released a vector network analysis module....

21/11/2025

The Best Movie Musicals on Every Streaming Platform

The Best Movie Musicals on Every Streaming Platform From Wicked to The Sound of Music, heres where to stream all the classic movie musicals and recent hits on...

20/11/2025

MLB Media-Rights Shakeup: NBC's New Three-Year Deal Covers Sunday Night Baseball,' Peacock-Exclusive Games, and More

MLB Media-Rights Shakeup: NBC's New Three-Year Deal Covers Sunday Night Bas...

20/11/2025

MLB Media-Rights Shakeup: New Deal Will Bring 30 National Games to ESPN's Linear Platform, MLB.TV Exclusively to ESPN App

MLB Media-Rights Shakeup: New Deal Will Bring 30 National Games to ESPN's Li...

20/11/2025

MLB Media-Rights Shakeup: Netflix Lands Opening Night, Home Run Derby, Field of Dreams

MLB Media-Rights Shakeup: Netflix Lands Opening Night, Home Run Derby, Field of ...

20/11/2025

MLB Media-Rights Shakeup Overview: ESPN, NBCU, Netflix Ink Three-Year Deals

MLB Media-Rights Shakeup Overview: ESPN, NBCU, Netflix Ink Three-Year DealsESPN gets new 30-game package, MLB.TV; NBC extends Sunday nights; Netflix adds tentpo...

20/11/2025

SVG Students To Watch: Henry Thuss, Indiana University

SVG Students To Watch: Henry Thuss, Indiana UniversityThe Southern California product has his goals set on the front benchBy Brandon Costa, Director of Digital ...

20/11/2025

Done+Dusted's Guy Carrington on Creating the Spectacular League of Legends World Championship Opening Ceremony

Done+Dusted's Guy Carrington on Creating the Spectacular League of Legends W...

20/11/2025

FIA Extreme H World Cup Host Broadcaster Aurora Goes Inside the Production of the Hydrogen-Fuelled Motorsport

FIA Extreme H World Cup Host Broadcaster Aurora Goes Inside the Production of th...

20/11/2025

Platinum White Paper: Amagi Utilizes Cloud Production for Sports Events - Multi-Camera Live Workflow with Remote Commentary and Graphics

Platinum White Paper: Amagi Utilizes Cloud Production for Sports Events - Multi-...

20/11/2025

2025 Sports Broadcasting Hall of Fame: Marc Herklotz, Steady Hand Behind the Scenes

2025 Sports Broadcasting Hall of Fame: Marc Herklotz, Steady Hand Behind the Sce...

20/11/2025

NFL Deep Dive: How 32 Cameras at Each Stadium Drive Virtual Measurement, Boundary Replays, and Skeletal Tracking

NFL Deep Dive: How 32 Cameras at Each Stadium Drive Virtual Measurement, Boundar...

20/11/2025

Zodiac Killer Project Refuses to Kill Its Darlings With an Avant-Garde Essay Film

Charlie Shackleton attends the 2025 Sundance Film Festival premiere of Zodiac K...

20/11/2025

Cracking the Code for Resilient, Secure Space-based Communications

L3Harris has achieved NSA Cybersecurity Directorate certification for its KSV-650 space hub end cryptographic unit, ensuring secure, adaptable communications fo...

20/11/2025

Continued Commitment and Investment in Poland

Left to Right: David Taubman, Regional Managing Director, Central and Eastern Europe; Arek Szalpuk, Poland Sr. Account Manager; Mr. Marcin Wi niewski, President...

20/11/2025

Oklahoma Community Television Overcomes ATSC 3.0 Translator Challenge

CINCINNATI Oklahoma Community Television (OCT) has selected GatesAir and Triveni Digital as key technology partners in an ATSC 3.0 deployment that establishes a...

20/11/2025

FCC Launches Wide-Ranging Examination of Network Affiliate Relations

WASHINGTON The Federal Communications Commission has opened a wide-ranging inquiry into the relations between broadcast networks and their affiliates that could...

20/11/2025

IAB: Creator Economy Ad Spend Now Dwarfs Ad Spend for Total Media Industry

NEW YORK Ad spend in the creator economy has more than doubled since 2021 from $13.9B to $29.5B in 2024 and that amount is projected to reach $37 billion in 202...

20/11/2025

Graduate Spotlight: Krysta Mirsik DePuy

Graduate Spotlight: Krysta Mirsik DePuy The educator from Hackettstown, New Jersey, shares how it took her 11 years to find the right graduate program for her...

20/11/2025

Rise Awards 2025 Celebrate the Industry's Best

LONDON The winners of the Rise Awards 2025, which recognize women and companies whose achievements have stood out in the media technology industry, have been an...

20/11/2025

Lionsgate, Debmar-Mercury Launch MovieSphere Gold Diginet

SANTA MONICA & NEW YORK Lionsgate's Worldwide Television Distribution Group and Debmar-Mercury have launched MovieSphere Gold in more than 30 million homes....