Sony Pixel Power calrec Sony

What Is Synthetic Data?

08/06/2021

Data is the new oil in today's age of AI, but only a lucky few are sitting on a gusher. So, many are making their own fuel, one that's both inexpensive and effective. It's called synthetic data.

What Is Synthetic Data? Synthetic data is annotated information that computer simulations or algorithms generate as an alternative to real-world data.

Put another way, synthetic data is created in digital worlds rather than collected from or measured in the real world.

It may be artificial, but synthetic data reflects real-world data, mathematically or statistically. Research demonstrates it can be as good or even better for training an AI model than data based on actual objects, events or people.

Users can generate synthetic data for autonomous vehicles using Python inside NVIDIA Omniverse. That's why developers of deep neural networks increasingly use synthetic data to train their models. Indeed, a 2019 survey of the field calls use of synthetic data one of the most promising general techniques on the rise in modern deep learning, especially computer vision that relies on unstructured data like images and video.

The 156-page report by Sergey I. Nikolenko of the Steklov Institute of Mathematics in St. Petersburg, Russia, cites 719 papers on synthetic data. Nikolenko concludes synthetic data is essential for further development of deep learning [and] many more potential use cases still remain to be discovered.

The rise of synthetic data comes as AI pioneer Andrew Ng is calling for a broad shift to a more data-centric approach to machine learning. He's rallying support for a benchmark or competition on data quality which many claim represents 80 percent of the work in AI.

Most benchmarks provide a fixed set of data and invite researchers to iterate on the code perhaps it's time to hold the code fixed and invite researchers to improve the data, he wrote in his newsletter, The Batch.

Augmented and Anonymized Versus Synthetic Data Most developers are already familiar with data augmentation, a technique that involves adding new data to an existing real-world dataset. For example, they might rotate or brighten an existing image to create a new one.

Given concerns and government policies about privacy, removing personal information from a dataset is an increasingly common practice. This is called data anonymization, and it's especially popular for text, a kind of structured data used in industries like finance and healthcare.

Augmented and anonymized data are not typically considered synthetic data. However, it's possible to create synthetic data using these techniques. For example, developers could blend two images of real-world cars to create a new synthetic image with two cars.

Why Is Synthetic Data So Important? Developers need large, carefully labeled datasets to train neural networks. More diverse training data generally makes for more accurate AI models.

The problem is gathering and labeling datasets that may contain a few thousand to tens of millions of elements is time consuming and often prohibitively expensive.

Enter synthetic data. A single image that could cost $6 from a labeling service can be artificially generated for six cents, estimates Paul Walborsky, who co-founded one of the first dedicated synthetic data services, AI.Reverie.

Cost savings are just the start. Synthetic data is key in dealing with privacy issues and reducing bias by ensuring you have the data diversity to represent the real world, Walborsky added.

Because synthetic datasets are automatically labeled and can deliberately include rare but crucial corner cases, it's sometimes better than real-world data.

What's the History of Synthetic Data? Synthetic data has been around in one form or another for decades. It's in computer games like flight simulators and scientific simulations of everything from atoms to galaxies.

Donald B. Rubin, a Harvard statistics professor, was helping branches of the U.S. government sort out issues such as an undercount especially of poor people in a census when he hit upon an idea. He described it in a 1993 paper often cited as the birth of synthetic data.

I used the term synthetic data in that paper referring to multiple simulated datasets, Rubin explained.

Each one looks like it could have been created by the same process that created the actual dataset, but none of the datasets reveal any real data - this has a tremendous advantage when studying personal, confidential datasets, he added.

In the wake of the Big Bang of AI, the ImageNet competition of 2012 when a neural network recognized objects faster than a human could, researchers started hunting in earnest for synthetic data.

Within a couple years, researchers were using rendered images in experiments, and it was paying off well enough that people started investing in products and tools to generate data with their 3D engines and content pipelines, said Gavriel State, a senior director of simulation technology and AI at NVIDIA.

Ford, BMW Generate Synthetic Data Banks, car makers, drones, factories, hospitals, retailers, robots and scientists use synthetic data today.

In a recent podcast, researchers from Ford described how they combine gaming engines and generative adversarial networks (GANs) to create synthetic data for AI training.

To optimize the process of how it makes cars, BMW created a virtual factory using NVIDIA Omniverse, a simulation platform that lets companies collaborate using multiple tools. The data BMW generates helps fine tune how assembly workers and robots work together to build cars efficiently.

Synthetic Data at the Hospital, Bank and Store Healthcare providers in fields such as medical imaging use synthetic data to train AI models while protecting patient privacy. For example, startup Curai trained a diagnostic model on 400,000 simul
LINK: https://blogs.nvidia.com/blog/2021/06/08/what-is-synthetic-data/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

23/05/2026

FOX Sports, IMS Productions Scale Up Indy 500 Production With New In-Car Cameras, AR Graphics, Cinematic Sets

In its second year as rightsholder, FOX Sports goes bigger across the board for ...

23/05/2026

Inside Apple TVs MLS iPhone Production with Royce Dickerson, Apple Live Sports, Executive Producer

Tonight's MLS matchup between the LA Galaxy and the Houston Dynamo FC will m...

23/05/2026

IK Multimedia reveal ReSing Voices Japanese Pack

AI-powered vocal tool gains first new language expansion IK Multimedia's AI-powered voice-creation software has seen a number of updates since it launch...

23/05/2026

Building a better future: Nielsen celebrates Global Volunteer Month and Earth Day 2026 with record participation

Nielsen Global Leadership Network graduates celebrate Earth Day 2026 Nielsen vo...

23/05/2026

Gray Media Names New Station General Managers

Share Copy link Facebook X Linkedin Bluesky Email...

23/05/2026

New CIMM Paper Urges Industry to Rethink How Media Is Evaluated

Share Copy link Facebook X Linkedin Bluesky Email...

23/05/2026

Lawo to Showcase Edge One, Efficient IP Workflows at InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

23/05/2026

Spectrum Launches Ultra-Low Latency Internet

Share Copy link Facebook X Linkedin Bluesky Email...

22/05/2026

Germanys Magenta TV Selects DMC to Provide FIFA World Cup Technical Support for Studios in Munich, New York City

Germany's Magenta TV, which will have 44 exclusive FIFA World Cup match broa...

22/05/2026

DAZN Grabs IFAF Flag Football Global Rights

DAZN, the world's leading sports entertainment platform, has acquired global broadcast rights to the International Federation of American Football's ( I...

22/05/2026

ATHLOS 2026 Season Set for October Debut in London; Aurora Media Worldwide Named Host Broadcast Partner

ATHLOS, the all-women's professional track and field league, has announced i...

22/05/2026

NATAS to Stream Sports, News, and Documentary Emmy Awards Live on YouTube

The National Academy of Television Arts & Sciences (NATAS) today announced that the 47th Annual Sports Emmy Awards and the 47th Annual News & Documentary Emmy A...

22/05/2026

Wooden Camera Rolls Out New Blackmagic URSA Accessories

Wooden Camera today announced the release of new accessories for the Blackmagic URSA Cine Immersive. The new lineup includes a redesigned Top Plate and Side Rai...

22/05/2026

YES Network, OTT Advisors Extend Streaming Partnership for Sixth Season

YES Network and OTT Advisors have announced a sixth consecutive season of their streaming partnership, continuing their collaboration on the Gotham app. OTT Adv...

22/05/2026

NESN Monster Week' Returns With Full Red Sox Broadcast From Atop Green Monster

NESN, New England's premier sports network, will again turn its camera to Fe...

22/05/2026

Dale Pro Audio RF Over Fiber Webinar Set for May 28

Dale Pro Audio is hosting an RF over Fiber Livestream Webinar on May 28 from 1-2:30 pm EST. With major sporting events and large-scale productions putting incre...

22/05/2026

Audio-Technica Appoints Humrichouser, Schanz to New Roles

Audio-Technica has announced key leadership appointments designed to further strengthen its sales organization and drive continued growth across the Americas. M...

22/05/2026

Scott Coker Launches Global MMA League With $60 Million in Backing

After nearly four decades shaping the global combat sports landscape, Scott Coker has announced a powerful return as he looks to build a new international mixed...

22/05/2026

Skyline Launches xOps Vanguard Runway for Autonomous Era

Skyline Communications, the company behind the globally deployed DataMiner xOps platform, today announced the launch of xOps Vanguard Runway, a strategic accele...

22/05/2026

The American Rodeo Takes Over Globe Life Field for Championship Weekend

For the fully onsite production, 30 cameras - including a SkyCam and Megalodon - will capture the action in Texas One of the world's biggest rodeo producti...

22/05/2026

Argentinas Torneos Taps Imagine Versio for Playout Operations Upgrade

Leading Argentina-based sports media company Torneos y Competencias S.A. has modernized its playout operations, implementing a fully redundant, multichannel env...

22/05/2026

Owl AI and Major League Pickleball Go Live with First-Ever AI Officiating System Powered by Broadcast Cameras and the Cloud

As the 2026 Major League Pickleball season kicks off this weekend in Dallas, it ...

22/05/2026

Shure, Edge Sound Research Look to Innovate via Partnership

Shure has become a minority investor in Edge Sound Research, a start-up company that is developing new experiential audio technologies that redefine how many au...

22/05/2026

SVG Rewind: MLBs UmpCam AR System Puts Fans Inside the Strike Zone Like Never Before

In advance of this year's Sports Emmy Awards, SVG is taking a deep dive into...

22/05/2026

Jelly Roll Offers Up 2026 Stanley Cup Playoff Theme Song for NHL, Amazon Music

The National Hockey League (NHL) and Amazon Music announced that GRAMMY Award-winning superstar Jelly Roll will provide the official theme song of the 2026 Stan...

22/05/2026

David Pogue, Andy Beach Keynotes Highlight Silicon Valley Video Summer Camp, July 14 at De Anza College

David Pogue will keynote SVV Summer Camp and discuss Apple at 50: How the World...

22/05/2026

FOX Sports, IMS Productions Scale Up Indy 500 Production in Year Two With New In-Car Cameras, AR Graphics, and Cinematic Sets

In its second year as rightsholder, FOX Sports goes bigger across the board for ...

22/05/2026

FOX Sports' Indy 500 Director Mitch Riggin on the Tech and Storytelling for the Greatest Spectacle in Racing

The broadcaster is drawing on lessons learned in its first year of covering the ...

22/05/2026

How Spotify's Rebuilt Ad Platform Is Delivering New Value for Brands

At our 2026 Investor Day, we shared an inside look at the rebuild of our advertising business. This pivot to our own purpose-built platform is already driving s...

22/05/2026

Spotify Levels Up Our Podcast Experience With New Features for Fans and Creators

Podcasting on Spotify continues to grow, and so do the ways listeners engage with it. At Investor Day 2026, we shared how we're building the next chapter of...

22/05/2026

CEDAR Audio introduce Icons Bundles

Limited-time collections now available Restoration experts CEDAR Audio have recently launched a new line of Icons plug-ins that make their powerful processo...

22/05/2026

Boss expand PS-1 Plugout Pedal

Three new classics join Model Pass line-up Boss' PX-1 Plugout Pedal offers an innovative approach to guitar pedals, providing users with a hardware stom...

22/05/2026

SGL Carbon commissions photovoltaic system and lays the foundation for a new nitrogen plant at its Meitingen site

At its Meitingen site, SGL Carbon has implemented two key projects to further de...

22/05/2026

Statement regarding 2026 National NAIDOC Lifetime Achievement Award for the late Rhoda Roberts AO

Statement regarding 2026 National NAIDOC Lifetime Achievement Award for the late...

22/05/2026

Polsat Reclaims Second Place and ByteDance Enters Top 10 as Polish Viewing Moves Beyond the Living Room in April

Latest data reveals steady distributor rankings, a seasonal shift toward digital...

22/05/2026

FCC Votes to Update Disaster Information Reporting System

Share Copy link Facebook X Linkedin Bluesky Email...

22/05/2026

Amagi delivers 30 per cent revenue growth in FY26 Adjuste...

Amagi Media Labs Limited (NSE: AMAGI, BSE: 544679), a cloud-native SaaS platform providing AI-enabled solutions to global media and entertainment companies, tod...

22/05/2026

Annima Post Relies on Cintel to Revive Classic Mexican Films

An nima Post Relies on Cintel to Revive Classic Mexican Films Brie Clayton May 22, 2026 0 Comments Film scanner and DaVinci Resolve Studio help manage...

22/05/2026

Boris FX Sapphire Adds Optical Beauty and Hypnotic Textures

Boris FX Sapphire Adds Optical Beauty and Hypnotic Textures Jessie Electa Petrov May 22, 2026 0 Comments The 2026.5 release introduces advanced defocu...

22/05/2026

Deployment Preserves Trusted Workflows While Enabling a P...

Deployment Preserves Trusted Workflows While Enabling a Path to UHD and SMPTE ST 2110 Leading Argentina-based sports media company Torneos y Competencias S.A....

22/05/2026

Study: AI Labeling Does Not Hurt Video Ad Performance

Share Copy link Facebook X Linkedin Bluesky Email...

22/05/2026

NAB Show Makes 200+ Sessions Available on Demand

Share Copy link Facebook X Linkedin Bluesky Email...

22/05/2026

Torneos Upgrades Multichannel Playout with Imagine's Versio

Share Copy link Facebook X Linkedin Bluesky Email...

22/05/2026

Ex-Husband, Current Husband, One Wild Rescue: Korean Action Comedy Husbands in Action' Premieres June 19

Back to All News Ex-Husband, Current Husband, One Wild Rescue: Korean Action Co...

22/05/2026

Beta Da Silva hosts live performances from 20 new Irish artists in Sessions from Oblivion on 2FM's New Music Show

Catch the latest in Irish music live from venues such as Whelan's, R is n Du...

21/05/2026

CBS Sports Expands WNBA Tip-Off Show To Cover Half of 20-Game, Regular-Season Package

Game Creek Video Columbia and Celtic, NEP Supershooter 8 will house onsite produ...