
A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.
While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.
Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.
For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.
This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.
A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.
Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.
Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.
A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.
The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.
An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.
Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.
Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.
Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.
For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.
With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.
Users Get Artistic Controls Several capabilities add to Fugatto's novelty.
During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.
The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.
I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.
In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.
The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.
Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.
A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.
The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.
Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.
One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.
They also scrutinized existing datasets to reveal new relationships among the dat
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
09/10/2026
September 10 2026, 06:00 (PDT) Dolby Expands Dolby OptiView Platform with New Capabilities at IBC 2026
New Sports Intelligence helps providers better unders...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
UK-based broadcast engineer Andrew Kemp has chosen a dhd.audio SX2 and XS3 as the basis of a transportable media-training resource. With a quarter-century of ex...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Cobalt Digital to Bring End-to-End ST 2110/IPMX Solutions and Practical Tools to NAB Show New York
Highlights include complete path from SDI to IP, C-band tran...
07/10/2026
Deerstalker Pictures The Exorcism of Nixie Shot With Blackmagic Cameras
Brie Clayton October 6, 2026
0 Comments
Fantasy series relied on Blackmagic ca...
07/10/2026
Nevion announces Providius NVRT integration with eMerge switches for expanded br...
07/10/2026
DriveShelf - A Filmmaker's Mac App for Finding Footage on Unplugged Drives
Brie Clayton October 6, 2026
0 Comments
DriveShelf for Mac catalogs eve...
07/10/2026
COW Jobs: Cinematographer, Video Editor - Baltimore, MD
Brie Clayton October 7, 2026
0 Comments
Cinematographer / Video Editor
September 23, 2026 ...
07/10/2026
Behind the Scenes at Berklee Homecoming 2026 Follow Marley Striem BM '25 behind the scenes as she manages a team of students and liaises with artists duri...
07/10/2026
Berklee College of Music and Berklee Valencia Named to Billboards 2026 List of T...
07/10/2026
Undercover filming shows add-on treatments with potential safety concerns on sal...
07/10/2026
Join Emma Doran, Justine Stafford, Killian Sundermann, Se n Burke and Michael Fry on RT 2 and RT Player tomorrow at 10.30pm
A host of guest stars include Siob...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
06/10/2026
A response to Mark Turner and the SVG AI Innovation Lab
SVG AI Innovation Lab Program Director Mark Turner's How Broadcast Sports Can Apply the Newsroom S...
06/10/2026
Cobalt Digital will exhibit at NAB Show New York 2026 (Booth 226, Javits Center, October 21-22), demonstrating its ST 2110/IPMX product line alongside solutions...
06/10/2026
LiveU and Airwise have announced a strategic partnership integrating LiveU's bonded video transmission with Airwise's drone fleet management, airspace a...
06/10/2026
Hotspur Labs, Tottenham Hotspur's corporate venture programme, will co-host ...
06/10/2026
SPORTEL Monaco 2026 will take place October 19-21 at the Grimaldi Forum in Monaco, with the conference programme running October 19-20.
Sessions include:
Mast...
06/10/2026
Zixi has announced the appointment of John Towers as Chief Financial Officer. Towers brings more than 20 years of experience in media and entertainment, includi...
06/10/2026
The Memphis Grizzlies, DAZN, and WATN/ABC24 have announced a partnership to simulcast 15 regular season Grizzlies games free over-the-air on ABC24 during the 20...
06/10/2026
The Minnesota Timberwolves, KARE 11, and DAZN have announced a partnership to simulcast 15 Timberwolves games free over-the-air on KARE 11 during the 2026-27 se...
06/10/2026
Sportel 2026 is set to be held in Monaco in less than two weeks (beginning Octob...
06/10/2026
The company's Pro Data portable storage unit enabled the Dallas Cowboys prod...
06/10/2026
Three new compact pedals announced
Warm Audio have just introduced a new guitar pedal line-up that offers compact, pedalboard-friendly versions of some of t...
06/10/2026
Drum-replacement tool gets major overhaul
WaveMachine Labs have announced that the latest version of their drum-replacement software is now available, and i...
06/10/2026
Firmware update implements user-requested features
Arturia's stage-going keyboard series has just received a significant update that addresses some of t...
06/10/2026
(4:51 PM) Wiesbaden, October 06, 2026. Based on business performance to date, pa...
06/10/2026
SBS invites audiences and communities to help shape the future of its language s...
06/10/2026
Manufacturer receives honors from both TV Tech and TVB Europe for advancements in compression technology
AMSTERDAM September 30, 2026 - Cobalt Digital has ...
06/10/2026
Cobalt Digital NAB NY Booth 226 // Journalists: Click to visit Cobalt
Highlight...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Imaginary Forces Paints a Portrait of Conflict in War
Brie Clayton October 6, 2026
0 Comments
Director Ronnie Koff of Imaginary Forces (IF) transforms...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Pebble, a leading specialist in automation, content management and integrated channel solutions, is supporting the continued modernisation of playout at Brazili...