
A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.
While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.
Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.
For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.
This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.
A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.
Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.
Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.
A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.
The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.
An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.
Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.
Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.
Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.
For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.
With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.
Users Get Artistic Controls Several capabilities add to Fugatto's novelty.
During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.
The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.
I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.
In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.
The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.
Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.
A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.
The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.
Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.
One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.
They also scrutinized existing datasets to reveal new relationships among the dat
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
22/07/2026
Among other tweaks to coverage, the broadcaster added announce positions in the ...
22/07/2026
The Australian Sports Technologies Network (ASTN) has reported a sharp acceleration in Australia's sportstech sector, with revenue growth rising from 10% in...
22/07/2026
From a retrofitted broadcast bus to packed Match Day Live! events, Bennett detai...
22/07/2026
TMRW Sports has named Brendan Donohue as President, Flag Football, where he will lead the company's professional flag football business in partnership with ...
22/07/2026
Stephen Arnold Music (SAM) designed and produced soundtracks for nine projection...
22/07/2026
Harmonic has announced a $25,000 donation to Global Empowerment Mission (GEM) fo...
22/07/2026
Telemundo has announced it has acquired from UC3 exclusive U.S. Spanish-language...
22/07/2026
The Rest Is Football, hosted by Gary Lineker, Alan Shearer, and Micah Richards, ...
22/07/2026
Sports Studio has announced a partnership with TheSoul Group to distribute Free ...
22/07/2026
Daktronics has concluded its 2026 High School Video Summit, a two-day hybrid event combining virtual sessions with hands-on host school visits across the countr...
22/07/2026
FOR-A America will exhibit at the MMCA Show in Pensacola, Florida (July 21-24), marking the company's first appearance at the house of worship AV event. FOR...
22/07/2026
The New York Yankees have selected Lumen Technologies (NYSE: LUMN) to connect Ya...
22/07/2026
Calrec will exhibit at IBC 2026 (Stand 8.C47) with demonstrations of its IP-native audio ecosystem built around the Argo console range, and will announce two ne...
22/07/2026
Solid State Logic has announced the appointment of Algam EKO as its distributor for SSL large format consoles in Italy, covering the Live, Broadcast, and Studio...
22/07/2026
FloSports has announced the launch of the FloSports Channel, a free ad-supported streaming channel premiering on Prime Video in the U.S. on July 21. The channel...
22/07/2026
Banana Ball TV (BTV), the broadcast operation of the Savannah Bananas, has insta...
22/07/2026
SiriusXM and Audacy have announced an agreement bringing 42 Audacy sports, news, and talk stations from 29 U.S. markets to SiriusXM. Twenty-three Audacy sports ...
22/07/2026
The collaboration produced more than 500 short-form videos across four cities th...
22/07/2026
At 2,600 sq. ft., the new home of digital content houses a main studio, podcast space, and an auxiliary control room
NFL franchises are accelerating their prod...
22/07/2026
Library now available in Engine Player
Another of Eduardo Tarilonte's popular sample libraries has just been moved over to Engine Audio's Engine Pla...
22/07/2026
All-in-one mic & DSP interface overhauled
Cloudvocal have just refreshed the design of their all-in-one mic, audio interface and DSP processor, introducing ...
22/07/2026
The final boss of the lost dual-op-amp circuit
Electro-Harmonix have just introduced the Deluxe Big Muff Pi 2, which they say offers their latest take on t...
22/07/2026
eds3_5_jq(document).ready(function($) { $(#eds_sliderM519).chameleonSlider_2_1({ content_source:......
22/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/07/2026
Proven managed IP solution helps broadcasters confidently navigate the future of broadcast distribution
LTN announces it has completed successful satellite-to-...
22/07/2026
Today, Chaos announced V-Ray 7, update 4, eliminating the need for asset conversions and tool switching, which slow down media and entertainment pipelines. Arti...
22/07/2026
An AI playground for the Project Indigo camera app
Marc Levoy July 22, 2026
0 Comments
At left is the user interface of our Playground, running on an ...
22/07/2026
Custom Consoles announces the completion of a large technical furniture project for Radio T l vision Suisse (RTS) at the networks recently inaugurated productio...
22/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
22/07/2026
Culver City, CA, July 21, 2026 3D streaming infrastructure provider Miris today announced glTF intake, including the binary .glb form, to its spatial streamin...
22/07/2026
Documentary Lo Invisible Captured with Blackmagic PYXIS 6K
Brie Clayton July 21, 2026
0 Comments
Camera's versatile design allowed filmmakers to c...
22/07/2026
AAF Bridge converts AAF and XML Timelines to Ableton Live - and back
Jerome Gardier July 21, 2026
0 Comments
A two-way converter for macOS and Windows...
22/07/2026
From Class to Concert: Explore a Student-Led Production of David Bowies Work Thr...
22/07/2026
A poisoned spy, a secret war, and a small English city contaminated by a deadly chemical weapon. Available on Sky Documentaries from 4 AugustWednesday 22 July 2...
22/07/2026
Audio at Parity: Comscore Launches Transcript-Level Targeting and Measurement wi...
22/07/2026
Arvato Systems and Lobster Enter into a Partnership and Combine the Integration Platform and Managed Services
Flexible integration solutions for cloud, hybrid...
22/07/2026
RT REPORTS NET SURPLUS OF 22.5 MILLION IN 2025...
22/07/2026
In the Opinion of the Censor airs Monday 27 July at 9.35pm on RT One and RT Player
In the Opinion of the Censor looks at how films for theatrical release wer...
22/07/2026
Before a healthcare robot can be useful in the real world, it has to learn how the physical world pushes back. Anatomy varies. Instruments bend, press, slip and...
21/07/2026
New research involving a cross-section of Mongolias independent media landscape finds enthusiasm for AI as a tool to improve efficiency, sustainability and resi...
21/07/2026
With winners Spain having already paraded the World Cup trophy through Madrid in front of millions of fans and the dust now settling on an epic tournament of th...