
A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.
While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.
Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.
For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.
This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.
A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.
Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.
Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.
A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.
The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.
An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.
Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.
Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.
Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.
For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.
With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.
Users Get Artistic Controls Several capabilities add to Fugatto's novelty.
During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.
The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.
I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.
In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.
The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.
Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.
A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.
The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.
Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.
One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.
They also scrutinized existing datasets to reveal new relationships among the dat
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
09/10/2026
September 10 2026, 06:00 (PDT) Dolby Expands Dolby OptiView Platform with New Capabilities at IBC 2026
New Sports Intelligence helps providers better unders...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
01/10/2026
Big Blue Marble has announced the addition of live multiview channels to its Clo...
01/10/2026
The company CEO sees AI as a tool in speeding and enhancing product development
...
01/10/2026
Production, tech, crewing, and logistics teams draw on experience of a busy inau...
01/10/2026
The Partnership Will Support Filmmakers Worldwide Through Original Short Films I...
01/10/2026
Third collaboration with celebrated composer
Spitfire Audio have once again teamed up with Grammy-nominated composer, producer and technologist BT, creating...
01/10/2026
FL Studio gets dedicated grid controller
Novation have teamed up with FL Studio developers Image-Line to create another dedicated controller for the popular...
01/10/2026
Six new USB interfaces introduced
Universal Audio's USB audio interface line-up has just been refreshed, gaining six new models aimed at musicians, prod...
01/10/2026
Siretta's SNYPER-5G cellular network analyser has been shortlisted for the IoT Global Awards 2026.
The IoT Global Awards recognise products, services and o...
01/10/2026
Johannesburg, South Africa, 30 September 2026 - The National Film and Video Foun...
01/10/2026
Mesh Broadcast Services Fits Big-Truck Audio Power into Compact Eclipse Unit with Calrec Argo M Canadian OB company Mesh Broadcast Services has fitted its Eclip...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
For more than two decades, Disguise has powered some of the most ambitious live events and experiences in the world from Bad Bunny and Coachella to the Burj K...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
The Australian football team can turn match footage into tactical insights to improve performance and maximize fan engagement
Vizrt, the AI platform for visual...
01/10/2026
When Mesh Broadcast Services set out to fit the pro audio capability of a large TV truck into a small audio room for high-profile sports events, Calrecs Argo M ...
01/10/2026
Emergent, the browser-native broadcast graphics platform, today announced the appointment of Tom Shelburne as Head of Business Development, North America, effec...
01/10/2026
Melbourne, Australia 1 October 2026: Mediaproxy, the global standard for software-based IP compliance monitoring and multiviewing solutions, will showcase maj...
01/10/2026
Clear-Com is supporting the education and development of future audiovisual industry professionals in Poland through the deployment of its Arcadia Central Sta...
01/10/2026
Telef nica Servicios Audiovisuales (TSA), the audiovisual services arm of global telecommunications group Telef nica, has supplied and commissioned an LED wall ...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
01/10/2026
Eleven minutes between headliners. The sponsor loop needs to be on screen now, and the only operator on hand is the one running lights. With the Green Hippo Est...
01/10/2026
DigiBox has announced a new partnership with FOR-A to represent the innovative MixBoard product range in the UK market, further expanding its portfolio of broa...
01/10/2026
Screen Australia-funded documentaries bring Australian experiences and new persp...
01/10/2026
Mavis Camera 8.2 adds iPhone 18 Pro support with variable aperture control
Brie Clayton September 30, 2026
0 Comments
Latest release also introduces P...
01/10/2026
Ghosts of the Gold Rush Shot with PYXIS 6K
Brie Clayton September 30, 2026
0 Comments
Blackmagic Design cameras expertly capture documentary while del...
01/10/2026
RE:Vision Effects - RSMB V7 is here!
Lori Freitag September 30, 2026
0 Comments
ReelSmart Motion Blur 7 is here, the RSMB V7 cycle is open. ReelSmart ...
01/10/2026
VideoGen raises $3.3M seed as its AI video platform passes 5 million users
Brie Clayton September 30, 2026
0 Comments
VideoGen co-founders Anton Koeni...
01/10/2026
New taproom, restaurant, lounge and theater in the historic Power Plant add to ATC's growing social hub
Fullsteam Brewery and CBC Real Estate celebrate th...
01/10/2026
Berklee and APM Music Launch Global Catalog Bringing Student-Created Music to Fi...
01/10/2026
X-Rite Pantone Launches Spectura to Advance Connected Color Quality from Lab to ...
01/10/2026
ITVX AND WARNER BROS. DISCOVERY ANNOUNCE NEW BRANDED CONTENT PARTNERSHIP
QUEST
TLC
Under a new branded content partnership between Warner Bros. Disco...
01/10/2026
Save $50 on Ivory 3 American Concert D and German DThis fall, bring the sound and expressive character of two world-class concert grands into your studio with s...
01/10/2026
Thursday 1 October 2026
Nominees revealed for 2026 Sky Arts Awards
Current slide, 0 0, undefined1
0
Irish National Opera and Jennifer Walshe: MARS
Jacob Al...
01/10/2026
Vital TV antenna replaced at SudburyHelicopter and broadcast engineers work together to upgrade TV services for 1.5M in the East
--
Share to Linkedin
Octobe...
01/10/2026
Arqiva appoints Simon Duffy as ChairSimon Duffy joins as Chair effective 01 October 2026
--
Share to Linkedin
October 1, 2026
Press Office
01 October 2026...
01/10/2026
Season 4 of the multi-award-winning hit entertainment series The 2 Johnnies Late Night Lock In is back, serving up another round of unpredictable laughs, unforg...
01/10/2026
New season begins Sunday, 4 October at 6pm on RT lyric fm and RT Listen
RT lyric fm's arts and culture documentary series The Lyric Feature returns on S...
01/10/2026
RT is proud to continue its commitment to supporting and celebrating the arts across Ireland this October, highlighting a vibrant programme of theatre, music, ...
01/10/2026
AI factories are built by the megawatt, even by the gigawatt. Each megawatt fact...
01/10/2026
Spooky season is streaming in. Alongside falling leaves, pumpkin spice and everything nice, 25 new games are joining GeForce NOW throughout October, including s...