Sony Pixel Power calrec Sony

Now Hear This: World's Most Flexible Sound Machine Debuts

25/11/2024

A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.

While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.

Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.

For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.

This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.

A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.

Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.

Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.

A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.

The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.

An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.

Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.

Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.

Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.

For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.

With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.

Users Get Artistic Controls Several capabilities add to Fugatto's novelty.

During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.

The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.

I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.

In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.

The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.

Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.

A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.

The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.

Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.

One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.

They also scrutinized existing datasets to reveal new relationships among the dat
LINK: https://blogs.nvidia.com/blog/fugatto-gen-ai-sound-model/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

23/04/2026

NAB Honors Rob Lowe and John Tesh With Hall of Fame Induction

Share Copy link Facebook X Linkedin Bluesky Email...

23/04/2026

Roku, Samsung Dominate CTV Platform Market in U.S.

Share Copy link Facebook X Linkedin Bluesky Email...

23/04/2026

G&D and VuWall Strengthen International Sales Team

Share Copy link Facebook X Linkedin Bluesky Email...

23/04/2026

The 2026 NAB Show Reports More than 58,000 Attendees

Share Copy link Facebook X Linkedin Bluesky Email...

23/04/2026

SmallHD Monitor Overlay License for Hi-5 and Hi-5 SX deli...

Partnership between ARRI and SmallHD brings new Hi-5 license Configurable monitor overlays adapt to individual working styles Supported by SmallHD monitors ru...

23/04/2026

Jeff Cronenweth ASC Sheds Light on Tron Ares with Astera

Lighting Master Cronenweth ASC brings a unique look to each grid world with the help of Astera Jeff Cronenweth on the set of Disney's TRON: ARES. Photo by...

23/04/2026

ZEISS Supreme Primes Shine in Star-Driven Short Dr Sam

DP Chloe Smolkin ( The Late Show, Kidz Bop ) joins director Danielle Beckmann and writer/actor Raji Ahsan behind the camera for the heartfelt short comedy Dr...

23/04/2026

Apply now to join the 2026 Producer Delegation to TIFF: The Market

Apply now to join the 2026 Producer Delegation to TIFF: The Market 23 April 2026 Screen Australia, in partnership with Ontario Creates, has opened application...

22/04/2026

Live From NAB 2026: Solid State Logics Berny Carpenter on Expanding System T With Virtual DSP, Cloud Workflows

Solid State Logic is advancing its System T platform with a stronger focus on IP...

22/04/2026

Live From NAB 2026: Dolbys Giles Baker on the Growth of Dolby OptiView, Immersive Vision and Audio for Live Sports

From immersive audio to live streaming, Dolby Laboratories is focused on the fut...

22/04/2026

Live From NAB 2026: Blackmagic Design's Bob Caniglia on Implementing Cinematic Looks in Live Broadcasts

Shallow depth-of-field cameras have taken the industry by storm. Its debut a han...

22/04/2026

NAB 2026: Eastern Kentucky University deploys campus-wide ST 2110 network with Riedel and Bridge Digital

Riedel Communications (Booth C4908) announced that Eastern Kentucky University (...

22/04/2026

SportsTechBuzz at NAB 2026, Day 4: Live Reports From the Show Floor in Vegas

The NAB Show is in full swing, and the SVG and SVG Europe editorial teams are chasing down the hottest stories from all over the Las Vegas Convention Center. He...

22/04/2026

NAB 2026: Blackmagic Design Announces URSA Cine 12K LF 100G

Blackmagic Design has announced the URSA Cine 12K LF 100G, a new model in the URSA Cine family adding 100G Ethernet for SMPTE 2110 live production output up to ...

22/04/2026

Live From NAB 2026: NEPs Martin Stewart Talks 40 Years, the NEP Platform, and Scaling for FIFA World Cup

Celebrating its 40th anniversary, NEP is leaning into hybrid production with the...

22/04/2026

Live From NAB 2026: NEPs Dan Murphy on NEP Platform, TFC, and the Shift to Software-Defined Workflows

NEP VP, Platform Dan Murphy sits down at the 2026 NAB Show to unpack what NEP P...

22/04/2026

Spotify and WNBA's New York Liberty Bring Basketball and Music Together With New Partnership

Spotify and the New York Liberty are teaming up to give music and basketball fan...

22/04/2026

The story of the Focusrite ISA preamp

New 20-minute documentary explores iconic design The Focusrite Room in Mesa, Arizona, where John Aquilino hosts the Studio Console 005. In 2025, Focusrite co...

22/04/2026

EverSync SP-10 wireless from Cloudvocal

Offers compact wireless solution for pedalboards Taiwanese audio brand Cloudvocal have announced the availability of a new pedalboard-friendly wireless syst...

22/04/2026

Arturia release Augmented Persia

Latest hybrid sampling/synthesis instrument arrives Arturia's Augmented series offerings rely on a mixture of sampling and synthesis, allowing users to ...

22/04/2026

Acustica Audio launch Salt 2

Combines three distinct analogue EQ emulations The latest addition to Acustica Audio's ever-expanding collection of analogue-emulation plug-ins combines...

22/04/2026

Analog Empire: Bass & Lead from Melda Production

Final instalment in vintage-inspired instrument series Analog Empire: Bass & Lead marks the final instalment in Melda Production's vintage hardware-insp...

22/04/2026

Strymon reveal the Canoga

Fuzz pedal joins all-analogue Series A line Given that Strymons reputation was built on unapologetically digital pedals, it was a little surprising to see t...

22/04/2026

SBS names shortlisted brands for 2026 SBS Media Sustainability Challenge

SBS names shortlisted brands for 2026 SBS Media Sustainability Challenge 22 April, 2026 Media releases National broadcaster also releases its second annual...

22/04/2026

The Frequency That Decides the Fight

Why Low Band Electronic Warfare Matters...

22/04/2026

Polish national football team play-off games top monthly programme list

The nation unites around football team's World Cup dream Warsaw, Poland, 20.04.26: Nielsen, a global leader in audience measurement, data, and media intell...

22/04/2026

Nielsen and the Polish Organisation of Advertisers announce strategic partnership to elevate marketing standards in Poland

Warsaw, Poland, 22.04.26: Nielsen, a global leader in audience measurement, data...

22/04/2026

Nielsen helps New Zealand brands expand internationally with greater clarity and confidence

New market intelligence offering gives businesses a clearer view of local consum...

22/04/2026

Glookast Unveils New UX, YouTube and Social Media Connectors, Premiere Panel, Cinnafilm Tachyon Plugin and More at NAB

Glookast Unveils New UX, YouTube and Social Media Connectors, Premiere Panel, Ci...

22/04/2026

Lightcraft Technology to Preview Spark Story at NAB 2026 with Interactive Previs Experience

Lightcraft Technology to Preview Spark Story at NAB 2026 with Interactive Previs...

22/04/2026

Bolin Demos New PTZ Cameras and Controller at 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

22/04/2026

Anchor Audio Launches Beacon 3

Share Copy link Facebook X Linkedin Bluesky Email...

22/04/2026

FCC Grants WSWB TV License Transfer to Sinclair

Share Copy link Facebook X Linkedin Bluesky Email...

22/04/2026

Telemundo Puerto Rico Streaming Channel Launches On Prime Video

Share Copy link Facebook X Linkedin Bluesky Email...

22/04/2026

Chyron Announces PRIME Translate

Share Copy link Facebook X Linkedin Bluesky Email...

22/04/2026

TV Tech Announces Winners of Best of Show Awards at 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

22/04/2026

VEON's Banglalink to Bring Starlink Mobile to Customers in Bangladesh

22 Apr 2026 VEON's Banglalink to Bring Starlink Mobile to Customers in Bangladesh Bangladesh becomes the third market where VEON and Starlink Mobile partne...

22/04/2026

FIRST LOOK FOR NEW U DRAMA SERIES HIT POINT

U have unveiled exclusive first-look images for their six-part police thriller Hit Point, starring Nick Blood (Day of the Jackal) and BAFTA nominee Saffron Hock...

22/04/2026

UKTV Highlights: Saturday May 9th -15th 2026

What can I watch on UKTV and stream on U this week? This week on UKTV and the free streaming service U, viewers can watch a range of new and returning programm...

22/04/2026

Sky announces fifth year of WNT Fund with 30,000 bursary supporting players and grassroots football

Wednesday 22 April 2026 Sky announces fifth year of WNT Fund with 30,000 bursa...

22/04/2026

This Earth Day, Discover the Sustainable Productions Behind Our Films and Series

Back to All News This Earth Day, Discover the Sustainable Productions Behind Our Films and Series Emma Stewart, Ph.D. Netflix Sustainability Officer Enterta...

22/04/2026

Retail Media Standards Are Expanding Into Commerce Media - Here's Why That Matters for Measurement

The move from Retail Media to Commerce Media is about broadening the scope of th...

22/04/2026

Dolby and BMW Bring Dolby Atmos to the BMW 7 Series, Expanding Immersive Audio Across Future Models

April 22 2026, 07:00 (PDT) Dolby and BMW Bring Dolby Atmos to the BMW 7 Series,...

22/04/2026

RT Licenses Stolen Sister to Pushkin

RT Documentary On One 7-part series breaks US market for first time RT Programme Sales has announced its first deal with a US distribution partner for its 7-...