Sony Pixel Power calrec Sony

Now Hear This: World's Most Flexible Sound Machine Debuts

25/11/2024

A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.

While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.

Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.

For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.

This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.

A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.

Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.

Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.

A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.

The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.

An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.

Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.

Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.

Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.

For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.

With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.

Users Get Artistic Controls Several capabilities add to Fugatto's novelty.

During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.

The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.

I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.

In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.

The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.

Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.

A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.

The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.

Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.

One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.

They also scrutinized existing datasets to reveal new relationships among the dat
LINK: https://blogs.nvidia.com/blog/fugatto-gen-ai-sound-model/...
See more stories from nvidia

Most recent headlines

09/11/2025

Dalet Unveils Agentic AI Media Workflows at IBC2025

Dalet today announced a transformative leap forward for media operations: Agentic Artificial Intelligence (AI) that unifies the Dalet ecosystem under one natura...

06/10/2025

France Tlvisions Wins Prestigious 2025 EBU Technology & Innovation Award in Groundbreaking Collaboration with Dalet

France T l visions, France's leading broadcaster, has received the 2025 EBU ...

16/09/2025

ENCO Introduces Raptor Cloud-Based Captioning For Live Streaming Video

NOVI, Mich. ENCO has unveiled Raptor, a cloud-based live streaming captioning encoder that injects the speed, power and reliability of real-time AI capabilities...

16/09/2025

Comcast NBCU and NBCU Local Award $2.5 Million to Nonprofits

NEW YORK Comcast NBCUniversal and NBCUniversal Local have announced that a total of $2.5 million has been awarded in 2025 to 69 nonprofit organizations servicin...

16/09/2025

Martin Euredjian Joins Atomos

AMSTERDAM Martin Euredjian has joined Atomos as director of advanced imaging and will lead innovation for advanced display technology....

16/09/2025

Calrec Unveils 48-Fader Argo M

AMSTERDAM Calrec introduced a 48-faced Argo M and showcased its largest Argo software updates at the recently concluded IBC2025....

16/09/2025

Lawo Unveils HOME Audio Shuffler App

AMSTERDAM Lawo introduced its HOME Audio Shuffler app, a replacement for a traditional baseband audio matrix within an IP-based Dynamic Media Facility, during t...

16/09/2025

77th Emmy Awards on CBS Deliver Largest Audience Since 2021

CBS is reporting that the 77TH Emmy Awards hosted by Nate Bargatze on Sunday Sept. 14 was seen by more than 7.42 million viewers on the CBS Television Network a...

15/09/2025

Get to Know This Fall's Filmmakers Through These 27 Sundance Institute-Supported Titles

Steve Zahn, Winona Ryder, Ethan Hawke, and Janeane Garofalo star in Ben Stiller&...

15/09/2025

aespa and Spotify Invite Fans to Unlock Their Inner Rich Man' With an Immersive MY VAULT Experience

Global K-Pop sensation aespa is redefining what it means to be rich with the r...

15/09/2025

Spotify's Free Experience Is Even Better-Here's How to Make the Most of It

Every day, millions of people around the world turn to Spotify to enjoy the audi...

15/09/2025

Brembo SGL Carbon Ceramic Brakes (BSCCB) successfully expands production capacity by 50% in Germany and Italy to meet rising demand

After months of intensive planning and implementation, Brembo SGL Carbon Ceramic...

15/09/2025

L3Harris Receives Multi-Year Javelin Solid Rocket Motor Contract

A U.S. Marine launches a Javelin shoulder-fired anti-tank missile during a training exercise. (Photo credit: U.S. Marine Corps)...

15/09/2025

TV Tech Unveils Best of Show Winners at IBC 2025

AMSTERDAM TV Tech has named its Best of Show Awards winners for IBC2025, which wraps up today. Entrants were judged by a panel of industry experts on the criter...

15/09/2025

SES SCORE Surpasses 600,000 of Transmission Hours, Delivering 900 Hours of Major Sports Content Daily

Unique sports content orchestration platform builds momentum among SES's cus...

15/09/2025

The Great Flood' Teaser Trailer Previews A Gripping Struggle for Survival

Back to All News The Great Flood' Teaser Trailer Previews A Gripping Struggle for Survival Entertainment 15 September 2025 GlobalSouth Korea Link copi...

15/09/2025

From Hardship to Hope: Typhoon Family' Presents a Tale of Youth, Crisis, and Resilience on October 11

Back to All News From Hardship to Hope: Typhoon Family' Presents a Tale of...

15/09/2025

Gloom Goes Global as Wednesday's' The Doom Tour Hits Five Continents for Season 2

Back to All News Gloom Goes Global as Wednesday's' The Doom Tour Hit...

15/09/2025

New study reveals overwhelming support for a more sustainable future

-- Opens door to growth in renewable energy New Delhi, India - 15th September -- Global business and industry leaders from around the world are joining technol...

14/09/2025

AMWA and EBU Form JT-DMF Joint Task Force on Dynamic Medi...

Partnership to address business and technical challenges of DMF adoption he Advanced Media Workflow Association (AMWA) and the European Broadcasting Union (EBU...

14/09/2025

Mantis' Trailer Previews High-Stakes Rivalries to Be No. 1 Contract Killer - Premieres September 26

Back to All News Mantis' Trailer Previews High-Stakes Rivalries to Be No. ...

13/09/2025

Cox Media Group's Misti Turnbull Inducted into the NATAS Silver Circle

ATLANTA Cox Media Group has announced that the company's vice president of news, Misty Turnbull has been inducted into the National Academy of Television Ar...

13/09/2025

Shotoku Debuts Swoop Cranes for Studio Robotics at IBC2025

AMSTERDAM Shotoku Broadcast Systems, a major developer of robotic systems, has announced plans to take studio robotics to the next level at IBC2025 by debuting ...

13/09/2025

Riedel Unveils Ultra-Light Bolero Mini Wireless Intercom...

At IBC2025 in Amsterdam, Riedel Communications unveiled Bolero Mini, the company's lightest and flattest wireless intercom beltpack to date. Designed to del...

13/09/2025

Shotoku Takes Studio Robotics to New Heights with IBC Deb...

Shotoku Broadcast Systems, the international developer of dependable, userfriendly robotic systems, is taking studio robotics to the next level at IBC 2025 with...

13/09/2025

The Bitmovin Video Developer Report 2025-26 Reveals Cost...

Bitmovin, a leading provider of video streaming solutions, today released the 9th annual Video Developer Report 2025/26, offering an in-depth look at the evolvi...

13/09/2025

Bitmovin and StreamShark Partner to Deliver High Quality...

Bitmovin, the leading provider of video streaming solutions, today announced a strategic partnership with StreamShark, the trusted video platform for enterprise...

13/09/2025

Ikegami Announces VFE-P711AD 7-inch OLED Multiformat On-C...

Ikegami has chosen IBC 2025 in Amsterdam as the launch venue for a major addition to its range of viewfinders. The new VFE-P711AD is a 7-inch high resolution OL...

13/09/2025

KitBash3D and Greyscalegorilla Announce Merger

Founder-led Merger to Fast Track R&D, Asset Library Upgrades, Tools and More; No Disruption to Pricing or Support for Users Today, KitBash3D, a pioneer in 3D a...

13/09/2025

Mavis Puts Itself at the Heart of Mobile Production

With NDI certification, Atomos integration, Grass Valley collaboration, and a new Monitor app, at this year's IBC, Mavis is showcasing a series of powerful...

13/09/2025

Creamsource Expands Vortex Family with Vortex24 Soft

Creamsource, maker of artisan LED lighting for film and television, has unveiled the Vortex24 Soft (V24S), a 1950W native soft light and the largest soft source...

13/09/2025

DAZN streams 2025 FIFA Club World Cup to billions of fans...

When international sports streaming service DAZN secured the global rights to the 2025 FIFA Club World Cup football tournament, it set out to deliver an unmatch...

13/09/2025

Riedel Communications Acquires hi human interface

Riedel Communications today announced the acquisition of hi human interface from Broadcast Solutions, bringing a powerful, vendor-agnostic control system to it...

13/09/2025

RTW chooses Calrec as technology partner for its AI ready...

Building on its long-term relationship with audio metering specialist RTW, Calrec has integrated the company's brand new TMxCore metering platform across it...

13/09/2025

Calrec unveils 48 fader Argo M at IBC2025 and demonstrate...

Calrec is expanding its family of future-ready self-contained Argo M control surfaces at IBC2025, with the addition of a brand new powerful 48-fader console. Co...

13/09/2025

Reaching Across the Isles: UK-LLM Brings AI to UK Languages With NVIDIA Nemotron

Celtic languages - including Cornish, Irish, Scottish Gaelic and Welsh - are the U.K.'s oldest living languages. To empower their speakers, the UK-LLM sover...

13/09/2025

SKY Perfect Modernizes Playout-to-Delivery with Harmonic

Harmonic's Software-Based XOS Advanced Media Processor Provides Unparalleled Efficiency and Unlocks New Business Models SAN JOSE, Calif. - Sept. 13, 2025 -...

13/09/2025

September 11, 2025

Researchers find brain region that fuels compulsive drinking Study by Scripps Research scientists shows how the brain learns to seek alcohol for relief, not jus...

12/09/2025

College Football Kickoff 2025: Fox Sports Ups Look as Canon, Sony Power Shallow Focus Coverage

College Football Kickoff 2025: Fox Sports Ups Look as Canon, Sony Power Shallow ...

12/09/2025

ABC/ESPN Excited For WNBA Postseason Coverage In Revamped Format

ABC/ESPN Excited For WNBA Postseason Coverage In Revamped FormatThe Finals moves to a best-of-seven series in 2025By Mark J Burns, SVG Contributor Friday, Sep...

12/09/2025

Rabbit Trap Pulsates With Folklore Dread

(L-R) Jade Croot, Rosy McEwen, and Bryn Chainey attend the 2025 Sundance Film Festival premiere of Rabbit Trap at Eccles Theatre on January 24, 2025, in Park ...

12/09/2025

Spotify's The Drop Weekly' Brings You the Week in New Releases, Straight From Our Editors

For fans, we know how important it is to stay plugged into music culture and dis...

12/09/2025

Agama and Consult Red announce RDK Accelerator integration

Link ping, Sweden and Shipley, United Kingdom, September 12, 2025 - Agama, the expert in video observability and analytics for service quality and customer expe...

12/09/2025

IBC2025 Opens for Business

IBC2025 began on Sept. 12, with exhibits and conferences running through Sept. 15 at the RAI Amsterdam Convention Center. Explore the full TV Tech coverage of t...

12/09/2025

The Best Fictional Bands (and the Artists Who Make Them Great)

The Best Fictional Bands (and the Artists Who Make Them Great) With Spinal Tap II: The End Continues hitting theaters and songs from KPop Demon Hunters ruling...

12/09/2025

Tom Baldassare Joins Advanced Systems Group

Industry veteran Tom Baldassare has joined Advanced Systems Group, LLC (ASG), a technology and services provider for media creatives and content owners, as a Se...

12/09/2025

Maxon Unveils a Brand New Look for its Growing Family of...

Maxon, maker of powerful, approachable software solutions for creators working in 2D and 3D design, motion graphics, visual effects, and more, today announced a...

12/09/2025

PlayBox Neo US Partners with AI-Media to Deliver Scalable...

PlayBox Neo, a leading provider of media playout solutions, has partnered with AI-Media, pioneering developers of AI-powered captioning technology, to integrate...