
A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.
While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.
Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.
For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.
This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.
A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.
Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.
Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.
A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.
The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.
An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.
Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.
Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.
Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.
For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.
With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.
Users Get Artistic Controls Several capabilities add to Fugatto's novelty.
During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.
The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.
I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.
In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.
The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.
Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.
A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.
The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.
Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.
One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.
They also scrutinized existing datasets to reveal new relationships among the dat
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
30/04/2026
The Professional Women's Hockey League (PWHL) concluded its third regular season on Saturday, reporting growth across attendance, viewership, digital engage...
30/04/2026
NBC Sports will air national MLB coverage on Sundays beginning May 3, with MLB Sunday Leadoff on Peacock and NBCSN at 12:30 p.m. ET, followed by the debut of th...
30/04/2026
Clear-Com has appointed Brian Grahn as Market Outreach Manager of the Americas and Ben Turnwell as Business Development Manager for EMEA live.
Grahn joined Cle...
30/04/2026
ARRI has introduced the cforce MAX, a new lens motor for the Hi-5 lens control system. The cforce MAX is twice as fast as the cforce plus motor it replaces whil...
30/04/2026
Knuerr, Voxtronic, and IHSE will jointly present an integrated control room solu...
30/04/2026
The CW Network and ESPN have announced an agreement to make the ESPN App the exclusive streaming home for all CW Sports live events. CW Sports will continue to ...
30/04/2026
Ed Sheeran's The Loop' tour launched in Auckland in January 2026 before moving on to Australia, with South America and the United States to follow late...
30/04/2026
Audinate has announced Dante Preset Creator, a free online tool for configuring Dante network settings before hardware is available on site. Presets created in ...
30/04/2026
Yahoo Sports has announced the appointment of Jarrod Schwarz as General Manager of Yahoo Sports. Schwarz will oversee product, design, and technology; revenue a...
30/04/2026
Nielsen has released a new report, Get Ready with Media Intelligence: 2026 FIFA World Cup Edition, examining U.S. soccer viewership trends, fan engagement, and ...
30/04/2026
USA Lacrosse and SportsEngine have announced an expanded partnership, naming Spo...
30/04/2026
Telos Alliance will participate in the 2026 Media Production and Technology Show (MPTS), taking place May 13-14 at Olympia London. Rather than exhibiting from a...
30/04/2026
The global streamer buys the U.S. DTC platform solutions provider for a reported...
30/04/2026
Tigo Sports, Paraguay's leading sports broadcaster, has upgraded its video infrastructure with Ateme solutions for live encoding, multiplexing, and signal c...
30/04/2026
World Rugby and IMG have announced a long-term media rights partnership focused on growing rugby in the United States ahead of the Men's and Women's Rug...
30/04/2026
For the second year in a row, Overtime and the National Women's Soccer League (NWSL) are teaming up through a renewed content partnership to bring fans even...
30/04/2026
The 22-year ESPN vet's responsibilities will reportedly be taken over by SVP Mike Foss...
30/04/2026
In-venue and creative video staffers at the professional and collegiate level ha...
30/04/2026
Amazon and Duke University have announced a multiyear agreement for Prime Video to present exclusive coverage of three Duke Blue Devils men's basketball neu...
30/04/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
30/04/2026
Music is evolving, and so are the ways you discover and connect with artists. In...
30/04/2026
Between April 22-29, the first inaugural Stockholm Music Week brought together thought leaders and partners across industries including music, tech, government,...
30/04/2026
Iconic large-format console upgraded
API's iconic Vision console has just been treated to an overhaul that aims to meet the demands of today's profe...
30/04/2026
Comes complete with miking accessories
The LCT 440 Pure has proven to be a popular member of Lewitt's mic line-up, offering impressive technical perform...
30/04/2026
24 October 2026 at The Octagon, Sheffield
Now in its eighth year, SynthFest UK is the largest event of its kind in the UK, bringing together the top keyboar...
30/04/2026
SBS & NITV LEAD NATIONAL RECONCILIATION WEEK 2026 WITH LANDMARK GULPILIL DOCUMEN...
30/04/2026
Rohde & Schwarz equips new Terminal 3 at Frankfurt Airport with security scanner...
30/04/2026
Rohde & Schwarz expands broadband amplifier portfolio with new power classes up ...
30/04/2026
Jennifer Ehle (Contagion, Zero Dark Thirty) and Alex Hassell (Rivals, Wasteman, ...
30/04/2026
MELBOURNE, Fla., April 29, 2026 - L3Harris Technologies (NYSE: LHX) today announ...
30/04/2026
MELBOURNE, Fla., April 30, 2026 - L3Harris Technologies (NYSE: LHX) reports first quarter 2026 results.
Highlights
Orders of $7.8 billion; book-to-bill of 1....
30/04/2026
Behind the Broadcast: The Sound of Elite Golf Golf is gaining popularity; the 2025 Ryder Cup achieved record-breaking viewing figures in the UK specifically, wi...
30/04/2026
STA VENERA, MALTA, APRIL 29, 2026 CPI Media, a voluntary organization within the Missionary Society of St Paul (MSSP) and a leading media house dedicated to p...
30/04/2026
New software platform delivers comprehensive timing measurement across production workflows...
30/04/2026
Once again, the UK Pavilion in Hall 5 of BroadcastAsia 2026 will feature the latest and best in technology specifically developed and tailored for modern media ...
30/04/2026
Avid powers faster workflows and next-generation immersive audio with latest Pro...
30/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
30/04/2026
Scalable broadcast-grade production over public internet, replacing traditional OB workflows...
30/04/2026
Live demonstrations highlight LCEVC ecosystem momentum, AI-powered video pipelines, and expansion across broadcast, streaming, and social media.
DTV (TV 3.0) ...
30/04/2026
Student Spotlight: Matthew Leon The dual major shares his path from community college to Berklee, and how his heritage influences his work.
April 29, 2026
B...
30/04/2026
Berklee Artists to Perform at Major Global Music Festivals As part of the Berklee Popular Music Institute, students will perform at Lollapalooza, Governors Ba...
30/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
30/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
30/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...