Sony Pixel Power calrec Sony

NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

14/06/2024

NVIDIA today announced Nemotron-4 340B, a family of open models that developers can use to generate synthetic data for training large language models (LLMs) for commercial applications across healthcare, finance, manufacturing, retail and every other industry.

High-quality training data plays a critical role in the performance, accuracy and quality of responses from a custom LLM - but robust datasets can be prohibitively expensive and difficult to access.

Through a uniquely permissive open model license, Nemotron-4 340B gives developers a free, scalable way to generate synthetic data that can help build powerful LLMs.

The Nemotron-4 340B family includes base, instruct and reward models that form a pipeline to generate synthetic data used for training and refining LLMs. The models are optimized to work with NVIDIA NeMo, an open-source framework for end-to-end model training, including data curation, customization and evaluation. They're also optimized for inference with the open-source NVIDIA TensorRT-LLM library.

Nemotron-4 340B can be downloaded now from the NVIDIA NGC catalog and from Hugging Face, where developers can also use the Train on DGX Cloud service to easily fine-tune open AI models. Developers will soon be able to access the models at ai.nvidia.com, where they'll be packaged as an NVIDIA NIM microservice with a standard application programming interface that can be deployed anywhere.

Navigating Nemotron to Generate Synthetic Data LLMs can help developers generate synthetic training data in scenarios where access to large, diverse labeled datasets is limited.

The Nemotron-4 340B Instruct model creates diverse synthetic data that mimics the characteristics of real-world data, helping improve data quality to increase the performance and robustness of custom LLMs across various domains.

Then, to boost the quality of the AI-generated data, developers can use the Nemotron-4 340B Reward model to filter for high-quality responses. Nemotron-4 340B Reward grades responses on five attributes: helpfulness, correctness, coherence, complexity and verbosity. It's currently first place on the Hugging Face RewardBench leaderboard, created by AI2, for evaluating the capabilities, safety and pitfalls of reward models.

In this synthetic data generation pipeline, (1) the Nemotron-4 340B Instruct model is first used to produce synthetic text-based output. An evaluator model, (2) Nemotron-4 340B Reward, then assesses this generated text - providing feedback that guides iterative improvements and ensures the synthetic data is accurate, relevant and aligned with specific requirements. Researchers can also create their own instruct or reward models by customizing the Nemotron-4 340B Base model using their proprietary data, combined with the included HelpSteer2 dataset.

Fine-Tuning With NeMo, Optimizing for Inference With TensorRT-LLM Using open-source NVIDIA NeMo and NVIDIA TensorRT-LLM, developers can optimize the efficiency of their instruct and reward models to generate synthetic data and to score responses.

All Nemotron-4 340B models are optimized with TensorRT-LLM to take advantage of tensor parallelism, a type of model parallelism in which individual weight matrices are split across multiple GPUs and servers, enabling efficient inference at scale.

Nemotron-4 340B Base, trained on 9 trillion tokens, can be customized using the NeMo framework to adapt to specific use cases or domains. This fine-tuning process benefits from extensive pretraining data and yields more accurate outputs for specific downstream tasks.

A variety of customization methods are available through the NeMo framework, including supervised fine-tuning and parameter-efficient fine-tuning methods such as low-rank adaptation, or LoRA.

To boost model quality, developers can align their models with NeMo Aligner and datasets annotated by Nemotron-4 340B Reward. Alignment is a key step in training LLMs, where a model's behavior is fine-tuned using algorithms like reinforcement learning from human feedback (RLHF) to ensure its outputs are safe, accurate, contextually appropriate and consistent with its intended goals.

Businesses seeking enterprise-grade support and security for production environments can also access NeMo and TensorRT-LLM through the cloud-native NVIDIA AI Enterprise software platform, which provides accelerated and efficient runtimes for generative AI foundation models.

Evaluating Model Security and Getting Started The Nemotron-4 340B Instruct model underwent extensive safety evaluation, including adversarial tests, and performed well across a wide range of risk indicators. Users should still perform careful evaluation of the model's outputs to ensure the synthetically generated data is suitable, safe and accurate for their use case.

For more information on model security and safety evaluation, read the model card.

Download Nemotron-4 340B models via NVIDIA NGC and Hugging Face. For more details, read the research papers on the model and dataset.

See notice regarding software product information.
LINK: https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm...
See more stories from nvidia

North America Stories

03/07/2026

Broadcast Solutions Acquires BFE

Share Copy link Facebook X Linkedin Bluesky Email...

03/07/2026

AdClarity Expands CTV Offering to 20 Countries

Share Copy link Facebook X Linkedin Bluesky Email...

03/07/2026

Mountain West Conference Launches SVOD Service With New Revenue Model

Share Copy link Facebook X Linkedin Bluesky Email...

03/07/2026

Death Star

Death Star Andy Marken July 2, 2026 0 Comments Look, I can't get involved. I've got work to do. It's not that I like the Empire; I hate i...

03/07/2026

Berklee's New Visual Identity: Honoring Our History, Building for What's Next

Berklee's New Visual Identity: Honoring Our History, Building for What's...

03/07/2026

July 01, 2026

Scripps Research scientists awarded $2M to advance global disease surveillance Two Gates Foundation grants will expand wastewater surveillance and AI-driven dis...

03/07/2026

July 02, 2026

Joan Pulupa joins Scripps Research faculty to study the organization of DNA in brain cells and its links to neurodegeneration Using smell-sensing neurons and ad...

02/07/2026

SVG Students To Watch: Abby Finke, University of Dayton

Entering her senior year, this hometown girl is paving a career in live sports production gaining experience in replay and audio and as a TD In the live-sports...

02/07/2026

SVG GameDay, Ep. 22: Winnipeg Jets Kyle Balharry - Going to Work in the Great White North

In-venue and creative video staffers at the professional and collegiate level ha...

02/07/2026

BLAST Reports $133 Million in 2025 Revenue, Opens New York Headquarters

BLAST, a competitive entertainment company focused on esports, has announced more than $133 million in revenue for 2025, representing more than 40% year-over-ye...

02/07/2026

Riedel and SKAARHOJ Expand Collaboration with SimplyLive Integration

Riedel Communications has announced official SKAARHOJ panel support for SimplyLive production workflows, enabled through the SimplyLive 2.1 release. The integra...

02/07/2026

Czech Fire Rescue Service Deploys LiveU Video Transmission for Emergency Operations

The Fire Rescue Service of the Czech Republic has deployed LiveU video-over-bond...

02/07/2026

Gravity Media USA Appoints Brittney Boston as Head of Business Development

Gravity Media USA has announced the appointment of Brittney Boston as Head of Business Development, effective July 1, 2026. Based in Nashville, Tennessee, Bosto...

02/07/2026

TwelveLabs Raises $100 Million in Series B Funding

TwelveLabs, a video intelligence company, has announced $100 million in Series B funding co-led by NEA and NAVER Ventures, with participation from Amazon, Radic...

02/07/2026

Pro Padel League Announces Broadcast Partnership With USA Sports for 2026 Season

The Pro Padel League (PPL) has announced a broadcast partnership with USA Sports that will air five PPL championship matches on CNBC during the 2026 season, the...

02/07/2026

LiveLike Powers Eight FIFA World Cup 2026 Fan Engagement Activations Across Five Continents

LiveLike, a digital fan engagement platform, has announced eight confirmed FIFA ...

02/07/2026

InfoComm 2026: Cobalt Digitals blueCORE Wins Futures Best of Show Award

Cobalt Digital has received Future's Best of Show Award, presented by AV Technology at InfoComm 2026, for its blueCORE family of standalone signal processor...

02/07/2026

Synamedia Appoints Dr. Tzvi Gerstl as CEO

Synamedia has announced the appointment of Dr. Tzvi Gerstl as Chief Executive Officer. Paul Segre, who has served as CEO for the past six years, will transition...

02/07/2026

Esports World Cup 2026 Announces Expanded Sony Partnership for Paris Event

The Esports Foundation (EF) and Sony Group Corporation have announced an expanded collaboration for the Esports World Cup 2026 (EWC), taking place in Paris, Fra...

02/07/2026

Zee Entertainment Secures Exclusive Bundesliga Rights in India for Five Years

Zee Entertainment Enterprises Ltd. ( Z') has announced exclusive broadcast and digital rights for the Bundesliga in India for five years, beginning with the...

02/07/2026

All Hands on Deck: NBCU Comes Together to Produce Ultra-Complex Sail4th 250 Broadcast on July 4

NBCU brings together News, Sports, Local, and Telemundo for a 50+ camera live pr...

02/07/2026

Release Rundown: What to Watch in July, From Gail Daughtry and the Celebrity Sex Pass to Murder 101

Zoey Deutch, John Slattery, Ken Marino, Miles Gutierrez-Riley, and Ben Wang appe...

02/07/2026

Imagine Communications Acquired by Lumine Group

Share Copy link Facebook X Linkedin Bluesky Email...

02/07/2026

Rise AV APAC Brings Mentoring Conversation to InfoComm As...

Following the successful launch of its inaugural APAC Mentoring Programme last month, the Rise AV APAC Regional Council will bring the conversation around mento...

02/07/2026

Blackmagic PYXIS 6K Used to Shoot Director Takahisa Zeze's Cry Out

Blackmagic PYXIS 6K Used to Shoot Director Takahisa Zeze's Cry Out Brie Clayton July 2, 2026 0 Comments Highly mobile camera supports tense and de...

02/07/2026

Broadcast Solutions acquires BFE, expanding its lead in European broadcast, media and communications infrastructure

Broadcast Solutions acquires BFE, expanding its lead in European broadcast, medi...

02/07/2026

Berklee Alum and Faculty Perform at Boston Public Library's 250th Anniversary Celebration of the Declaration of Independence

Berklee Alum and Faculty Perform at Boston Public Library's 250th Anniversar...

02/07/2026

Broadcast Solutions acquires BFE

Broadcast Solutions GmbH, a leading systems integrator and provider of innovative solutions for the broadcast media industry, is acquiring BFE Studio und Medien...

02/07/2026

LiveMode builds agile content ingest with Cinegy

Cinegy GmbH, the premier provider of software-defined television technology, has extended the ingest facility at leading Brazilian sports company LiveMode, work...

02/07/2026

Synamedia Appoints Dr Tzvi Gerstl CEO

Share Copy link Facebook X Linkedin Bluesky Email...

02/07/2026

Cobalt Digitals blueCORE Wins Futures Best of Show Award...

Standalone processors acknowledged for the innovation and value they bring to Pro AV Cobalt Digital, a leading designer and manufacturer of signal processing ...

02/07/2026

Synamedia Appoints Dr Tzvi Gerstl as CEO as Company Enter...

Synamedia announced today the appointment of Dr Tzvi Gerstl as Chief Executive Officer. Paul Segre, who has served as CEO for the past six years, will transitio...

02/07/2026

Maxon Autograph: Introduction to working with Tables

Maxon Autograph: Introduction to working with Tables Simon Ubsdell July 1, 2026 0 Comments An overview of Autograph's ridiculously powerful tables...

02/07/2026

Boston Conservatory's Soire Breaks Records to Fund Student Scholarships

Boston Conservatory's Soir e Breaks Records to Fund Student Scholarships The event achieved 127 percent of its fundraising goal in an evening celebrating ...

02/07/2026

How Adam Rosenwach Pivoted from Music to Med Tech Without Missing a Beat

How Adam Rosenwach Pivoted from Music to Med Tech Without Missing a Beat What do the rehearsal room and the boardroom have in common? More than you might thin...

02/07/2026

Warner Bros. Discovery UK & Ireland backs Unacceptable for a second series on TLC ahead of Sunday's premiere

Warner Bros. Discovery UK & Ireland backs Unacceptable for a second series on TL...

02/07/2026

Joyride Through July With 12 Games Coming to GeForce NOW

Summer is heating up - and GeForce NOW is taking players along for the ride. Start the month with Monopoly: Star Wars Heroes vs. Villains, bringing a galaxy fa...

01/07/2026

Broadcast Management Group Appoints Kathy Samuels as Director of Creative Services

Broadcast Management Group (BMG) has announced the appointment of Kathy Samuels ...

01/07/2026

Shade Launches Custom Objects and Automations

Shade has announced Custom Objects and Automations, a platform expansion releasing June 29, 2026, that adds database and workflow automation capabilities direct...

01/07/2026

FOR-A America Adds Two Regional Sales Leaders

FOR-A America has announced the addition of Jaz Wray and Fernando Cruz to its U.S. sales team. Both report to Ernie Leon, Senior VP and Head of Sales and Strate...

01/07/2026

NBC Sports To Present All 15 MLB Games Nationally on July 4 Weekend Star-Spangled Sunday'

NBC Sports will air all 15 MLB games nationally on Sunday, July 5, across NBC, P...

01/07/2026

Clear-Com Upgrades Wireless Communications for Jeopardy! and Wheel of Fortune

Clear-Com has announced a wireless communications upgrade for Jeopardy! and Wheel of Fortune, deploying FreeSpeak II and FreeSpeak Icon systems across both prod...

01/07/2026

England Deploys Sony STATSports Live GPS Tracking at FIFA World Cup 2026

England's performance team will use Sony's STATSports APEX GPS tracking system to monitor player physical data in real time during FIFA World Cup 2026 m...

01/07/2026

Adder Technology Appoints Neil Hillier as CEO

Adder Technology has announced the appointment of Neil Hillier as Chief Executive Officer, effective July 1, 2026. Hillier succeeds Adrian Dickens, who transiti...

01/07/2026

Bitcentral Splits Into Two Companies: Bitcentral and ViewNexa

Bitcentral, Inc. has announced a strategic transaction creating two separate companies. The Production and Playout business will continue as Bitcentral, now own...

01/07/2026

DAZN48 Creator Initiative Draws Global Participation for FIFA World Cup 2026

DAZN has announced results from DAZN48, its creator initiative for the FIFA World Cup 2026. Launched in April 2026, the program received thousands of applicatio...

01/07/2026

IDEA To Induct Daktronics Sarah Rose Into Hall of Fame

Sarah Rose, VP, global services, Daktronics (NASDAQ: DAKT), will be inducted into the Information Display and Entertainment Association (IDEA) Hall of Fame at t...

01/07/2026

Gravity Media Delivers Distribution and Streaming Services for World Economic Forum in Dalian

Gravity Media and the World Economic Forum's production team provided broadc...

01/07/2026

Insight Productions Launches Insight Storm, a 53-Foot Esports Broadcast Truck

Insight Productions has announced the launch of Insight Storm, a 53-foot mobile broadcast unit built for esports production. The truck is built around a Ross Vi...

01/07/2026

ESPN Announces America 250 Content Initiatives Across Platforms

ESPN has announced several content initiatives marking America's 250th anniversary, as part of The Walt Disney Company's Disney Celebrates America pro...