
Data is the new oil in today's age of AI, but only a lucky few are sitting on a gusher. So, many are making their own fuel, one that's both inexpensive and effective. It's called synthetic data.
What Is Synthetic Data? Synthetic data is annotated information that computer simulations or algorithms generate as an alternative to real-world data.
Put another way, synthetic data is created in digital worlds rather than collected from or measured in the real world.
It may be artificial, but synthetic data reflects real-world data, mathematically or statistically. Research demonstrates it can be as good or even better for training an AI model than data based on actual objects, events or people.
Users can generate synthetic data for autonomous vehicles using Python inside NVIDIA Omniverse. That's why developers of deep neural networks increasingly use synthetic data to train their models. Indeed, a 2019 survey of the field calls use of synthetic data one of the most promising general techniques on the rise in modern deep learning, especially computer vision that relies on unstructured data like images and video.
The 156-page report by Sergey I. Nikolenko of the Steklov Institute of Mathematics in St. Petersburg, Russia, cites 719 papers on synthetic data. Nikolenko concludes synthetic data is essential for further development of deep learning [and] many more potential use cases still remain to be discovered.
The rise of synthetic data comes as AI pioneer Andrew Ng is calling for a broad shift to a more data-centric approach to machine learning. He's rallying support for a benchmark or competition on data quality which many claim represents 80 percent of the work in AI.
Most benchmarks provide a fixed set of data and invite researchers to iterate on the code perhaps it's time to hold the code fixed and invite researchers to improve the data, he wrote in his newsletter, The Batch.
Augmented and Anonymized Versus Synthetic Data Most developers are already familiar with data augmentation, a technique that involves adding new data to an existing real-world dataset. For example, they might rotate or brighten an existing image to create a new one.
Given concerns and government policies about privacy, removing personal information from a dataset is an increasingly common practice. This is called data anonymization, and it's especially popular for text, a kind of structured data used in industries like finance and healthcare.
Augmented and anonymized data are not typically considered synthetic data. However, it's possible to create synthetic data using these techniques. For example, developers could blend two images of real-world cars to create a new synthetic image with two cars.
Why Is Synthetic Data So Important? Developers need large, carefully labeled datasets to train neural networks. More diverse training data generally makes for more accurate AI models.
The problem is gathering and labeling datasets that may contain a few thousand to tens of millions of elements is time consuming and often prohibitively expensive.
Enter synthetic data. A single image that could cost $6 from a labeling service can be artificially generated for six cents, estimates Paul Walborsky, who co-founded one of the first dedicated synthetic data services, AI.Reverie.
Cost savings are just the start. Synthetic data is key in dealing with privacy issues and reducing bias by ensuring you have the data diversity to represent the real world, Walborsky added.
Because synthetic datasets are automatically labeled and can deliberately include rare but crucial corner cases, it's sometimes better than real-world data.
What's the History of Synthetic Data? Synthetic data has been around in one form or another for decades. It's in computer games like flight simulators and scientific simulations of everything from atoms to galaxies.
Donald B. Rubin, a Harvard statistics professor, was helping branches of the U.S. government sort out issues such as an undercount especially of poor people in a census when he hit upon an idea. He described it in a 1993 paper often cited as the birth of synthetic data.
I used the term synthetic data in that paper referring to multiple simulated datasets, Rubin explained.
Each one looks like it could have been created by the same process that created the actual dataset, but none of the datasets reveal any real data - this has a tremendous advantage when studying personal, confidential datasets, he added.
In the wake of the Big Bang of AI, the ImageNet competition of 2012 when a neural network recognized objects faster than a human could, researchers started hunting in earnest for synthetic data.
Within a couple years, researchers were using rendered images in experiments, and it was paying off well enough that people started investing in products and tools to generate data with their 3D engines and content pipelines, said Gavriel State, a senior director of simulation technology and AI at NVIDIA.
Ford, BMW Generate Synthetic Data Banks, car makers, drones, factories, hospitals, retailers, robots and scientists use synthetic data today.
In a recent podcast, researchers from Ford described how they combine gaming engines and generative adversarial networks (GANs) to create synthetic data for AI training.
To optimize the process of how it makes cars, BMW created a virtual factory using NVIDIA Omniverse, a simulation platform that lets companies collaborate using multiple tools. The data BMW generates helps fine tune how assembly workers and robots work together to build cars efficiently.
Synthetic Data at the Hospital, Bank and Store Healthcare providers in fields such as medical imaging use synthetic data to train AI models while protecting patient privacy. For example, startup Curai trained a diagnostic model on 400,000 simul
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
23/05/2026
In its second year as rightsholder, FOX Sports goes bigger across the board for ...
23/05/2026
Tonight's MLS matchup between the LA Galaxy and the Houston Dynamo FC will m...
23/05/2026
AI-powered vocal tool gains first new language expansion
IK Multimedia's AI-powered voice-creation software has seen a number of updates since it launch...
23/05/2026
Nielsen Global Leadership Network graduates celebrate Earth Day 2026
Nielsen vo...
23/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
23/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
23/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
23/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/05/2026
Germany's Magenta TV, which will have 44 exclusive FIFA World Cup match broa...
22/05/2026
DAZN, the world's leading sports entertainment platform, has acquired global broadcast rights to the International Federation of American Football's ( I...
22/05/2026
ATHLOS, the all-women's professional track and field league, has announced i...
22/05/2026
The National Academy of Television Arts & Sciences (NATAS) today announced that the 47th Annual Sports Emmy Awards and the 47th Annual News & Documentary Emmy A...
22/05/2026
Wooden Camera today announced the release of new accessories for the Blackmagic URSA Cine Immersive. The new lineup includes a redesigned Top Plate and Side Rai...
22/05/2026
YES Network and OTT Advisors have announced a sixth consecutive season of their streaming partnership, continuing their collaboration on the Gotham app. OTT Adv...
22/05/2026
NESN, New England's premier sports network, will again turn its camera to Fe...
22/05/2026
Dale Pro Audio is hosting an RF over Fiber Livestream Webinar on May 28 from 1-2:30 pm EST. With major sporting events and large-scale productions putting incre...
22/05/2026
Audio-Technica has announced key leadership appointments designed to further strengthen its sales organization and drive continued growth across the Americas. M...
22/05/2026
After nearly four decades shaping the global combat sports landscape, Scott Coker has announced a powerful return as he looks to build a new international mixed...
22/05/2026
Skyline Communications, the company behind the globally deployed DataMiner xOps platform, today announced the launch of xOps Vanguard Runway, a strategic accele...
22/05/2026
For the fully onsite production, 30 cameras - including a SkyCam and Megalodon - will capture the action in Texas
One of the world's biggest rodeo producti...
22/05/2026
Leading Argentina-based sports media company Torneos y Competencias S.A. has modernized its playout operations, implementing a fully redundant, multichannel env...
22/05/2026
As the 2026 Major League Pickleball season kicks off this weekend in Dallas, it ...
22/05/2026
Shure has become a minority investor in Edge Sound Research, a start-up company that is developing new experiential audio technologies that redefine how many au...
22/05/2026
In advance of this year's Sports Emmy Awards, SVG is taking a deep dive into...
22/05/2026
The National Hockey League (NHL) and Amazon Music announced that GRAMMY Award-winning superstar Jelly Roll will provide the official theme song of the 2026 Stan...
22/05/2026
David Pogue will keynote SVV Summer Camp and discuss Apple at 50: How the World...
22/05/2026
In its second year as rightsholder, FOX Sports goes bigger across the board for ...
22/05/2026
The broadcaster is drawing on lessons learned in its first year of covering the ...
22/05/2026
At our 2026 Investor Day, we shared an inside look at the rebuild of our advertising business. This pivot to our own purpose-built platform is already driving s...
22/05/2026
Podcasting on Spotify continues to grow, and so do the ways listeners engage with it. At Investor Day 2026, we shared how we're building the next chapter of...
22/05/2026
Limited-time collections now available
Restoration experts CEDAR Audio have recently launched a new line of Icons plug-ins that make their powerful processo...
22/05/2026
Three new classics join Model Pass line-up
Boss' PX-1 Plugout Pedal offers an innovative approach to guitar pedals, providing users with a hardware stom...
22/05/2026
At its Meitingen site, SGL Carbon has implemented two key projects to further de...
22/05/2026
Statement regarding 2026 National NAIDOC Lifetime Achievement Award for the late...
22/05/2026
Latest data reveals steady distributor rankings, a seasonal shift toward digital...
22/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/05/2026
Amagi Media Labs Limited (NSE: AMAGI, BSE: 544679), a cloud-native SaaS platform providing AI-enabled solutions to global media and entertainment companies, tod...
22/05/2026
An nima Post Relies on Cintel to Revive Classic Mexican Films
Brie Clayton May 22, 2026
0 Comments
Film scanner and DaVinci Resolve Studio help manage...
22/05/2026
Boris FX Sapphire Adds Optical Beauty and Hypnotic Textures
Jessie Electa Petrov May 22, 2026
0 Comments
The 2026.5 release introduces advanced defocu...
22/05/2026
Deployment Preserves Trusted Workflows While Enabling a Path to UHD and SMPTE ST 2110
Leading Argentina-based sports media company Torneos y Competencias S.A....
22/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/05/2026
Back to All News
Ex-Husband, Current Husband, One Wild Rescue: Korean Action Co...
22/05/2026
Catch the latest in Irish music live from venues such as Whelan's, R is n Du...
21/05/2026
Game Creek Video Columbia and Celtic, NEP Supershooter 8 will house onsite produ...