Sony Pixel Power calrec Sony

Now Hear This: World's Most Flexible Sound Machine Debuts

25/11/2024

A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.

While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.

Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.

For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.

This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.

A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.

Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.

Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.

A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.

The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.

An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.

Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.

Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.

Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.

For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.

With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.

Users Get Artistic Controls Several capabilities add to Fugatto's novelty.

During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.

The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.

I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.

In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.

The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.

Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.

A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.

The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.

Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.

One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.

They also scrutinized existing datasets to reveal new relationships among the dat
LINK: https://blogs.nvidia.com/blog/fugatto-gen-ai-sound-model/...
See more stories from nvidia

Most recent headlines

26/11/2025

Fubo, NBCUniversal Trade Barbs in Carriage Dispute

Following the Nov. 21 blackout of NBCUniversal channels on Fubo, the two sides have traded barbs about their inability to reach a new carriage deal....

26/11/2025

Global Sports Rights Spending to Top $78 Billion in 2030

LONDON As TV sports rights become increasingly important for both broadcasters and streamers, Ampere Analysis predicts global investment in the genre will surpa...

26/11/2025

Vubiquity Earns AWS Media & Entertainment Competency Status

LOS ANGELES Vubiquity said it has achieved the Amazon Web Services (AWS) Media & Entertainment Competency as part of the AWS Partner Network (APN). This designa...

26/11/2025

Comcast Pays $1.5 Million to Settle FCC Data Breach Probe

WASHINGTON The Federal Communications Commission's Enforcement Bureau said it has entered into a consent decree with Comcast calling for the cable company t...

26/11/2025

Berklee Named to the Hollywood Reporters Top Music Schools List

Berklee Named to the Hollywood Reporters Top Music Schools List The publication highlights the college's screen scoring program, industry partnerships, and ...

25/11/2025

Tracy Bonareri Onchoke: Winner, Young Journalist Award 2025

Tracy Bonareri Onchoke, an investigative journalist from Kenya is the winner of the Thomson Foundation's Young Journalist Award 2025. The 26-year-old-sele...

25/11/2025

SVG All-Stars: Blayke Scheer, Senior Director, Creative Content, YES Network

SVG All-Stars: Blayke Scheer, Senior Director, Creative Content, YES NetworkThe Indiana alum has turned storytelling into an artform for more than two decadesBy...

25/11/2025

Op-Ed: With FCC's C-Band Auction on the Horizon, Broadcasters Need Proven, Cost-Effective Alternatives

Op-Ed: With FCC's C-Band Auction on the Horizon, Broadcasters Need Proven, C...

25/11/2025

Analysis: Is Baller League Really the Future of Sport?

Analysis: Is Baller League really the future of sport? By Callum McCarthy, Editor-at-Large Tuesday, November 25, 2025 - 10:10 Print This Story With KSI on...

25/11/2025

Platinum Whitepaper: The Growth of Broadcast in the World of Major Large Scale Events with SOS Global

Platinum Whitepaper: The Growth of Broadcast in the World of Major Large Scale E...

25/11/2025

SVG Summit 2025 Preview: SVG Women's Sports Workshop

SVG Summit 2025 Preview: SVG Women's Sports WorkshopBy Samantha Gabay Tuesday, November 25, 2025 - 10:27 am Print This Story | Subscribe Story Highlig...

25/11/2025

SVG New Sponsor Spotlight: CacheFly's Matt Levine on the Evolving Role of the CDN and Prioritizing Throughput

SVG New Sponsor Spotlight: CacheFly's Matt Levine on the Evolving Role of th...

25/11/2025

Peacock's EA SPORTS Madden NFL Cast Levels Up on Thanksgiving With SkyCam as the Primary Angle and More Madden Elements

Peacock's EA SPORTS Madden NFL Cast Levels Up on Thanksgiving With SkyCam as...

25/11/2025

Sauna Is an Intimate Exploration of Queer Love and Identity

Mathias Broe attends the 2025 Sundance Film Festival premiere of Sauna at Library Center Theatre. (Photo by Michael Hurcomb/Shutterstock for Sundance Film Fes...

25/11/2025

5 Reasons to Try Spotify Premium This Holiday Season

The best playlists, podcasts, and audiobooks bring a little extra magic to your daily routine. With new features and offerings, Spotify Premium delivers even mo...

25/11/2025

New Study Reveals Australians Love Discovering New Music

Comprehensive new research confirms what we already knew: Australian music fans love the quality, quantity, and access they have to new and local music on strea...

25/11/2025

Why Use a SIM Card With The SNYPER-5G

Applicable Products Objectives The purpose of this application note is to give a brief background on 5G (NR) wireless communication an explain the reason a SN...

25/11/2025

Lionsgate and Nielsen expand partnership to deliver first-ever combined FAST channel and digital network measurement

Nielsen will now measure both Lionsgate's FAST channel MovieSphere and Movie...

25/11/2025

AP Switches to DaVinci Resolve Studio for Global News Production

FREMONT, Calif. Blackmagic Design said the Associated Press has completed the transition of its global video-editing platform to DaVinci Resolve Studio....

25/11/2025

Berklees Inaugural Nat King Cole and Natalie Cole Scholarship Awarded to Paris Pineyro

Berklees Inaugural Nat King Cole and Natalie Cole Scholarship Awarded to Paris P...

25/11/2025

Traditional TV Players Gained Viewers in October: Nielsen Gauge

NEW YORK NFL and college football coverage, the MLB postseason and the new fall broadcast-TV season contributed to major gains for traditional media companies a...

25/11/2025

Tower Products CEO Jim Veltrie to Retire Dec. 30

SAUGERTIES, N.Y. Tower Products, a manufacturer and distributor of pro video and audio equipment here, said President and CEO Jim Veltrie will retire from the c...

25/11/2025

Sinclair Makes Unsolicited Bid to Buy Scripps at $7 a Share

Following last week's disclosure that it had acquired a 8.2% stake in E.W. Scripps, Sinclair has filed papers with the Securities and Exchange Commission pr...

25/11/2025

VEON's QazCode and MeetKai Sign Agreement to Power National LLM Training and Local-Language Agentic Services Across VEON Markets

25 Nov 2025 VEON's QazCode and MeetKai Sign Agreement to Power National LLM...

25/11/2025

UKTV acquires three shows from Paramount Global Content Distribution for U, U&W and U&alibi

UKTV has acquired a high-profile slate of US dramas from Paramount Global Conten...

25/11/2025

Will Sharpe, Paul Bettany and Gabrielle Creevy star in a spectacular five-part event series Amadeus: Full Trailer Released

A symphony of genius, rivalry and vengeance, boldly reimagined from Peter Shaffe...

25/11/2025

Bradford Young named 2025 FilmLight Colour Awards Jury President'

Article courtesy of Cinematography World Read the article FilmLight has finalised the prestigious 2025 FilmLight Colour Awards jury and welcomed award-winning...

25/11/2025

Correccin de color en Chespirito: Sin Querer Queriendo

Article courtesy of Prensario Read the article La serie fue dirigida por Juli n de Tavira, Rodrigo Santos, y David Leche Ruiz, con direcci n de fotograf a a...

25/11/2025

Nosferatu,' Sinners,' The Studio' and Severance' Colourists Nominated for FilmLight Colour Awards

Article courtesy of The Hollywood Reporter Read the article The awards, celebr...

25/11/2025

Harbor rolls out Nara globally

Article courtesy of Televisual Read the article Already live in Los Angeles and rolling out in New York and London, Nara gives producers, colourists, conform ...

25/11/2025

ARTONE FILM integrates Baselight M

Article courtesy of Digital Media World Read the article ARTONE post-house in Tokyo is the first facility in Japan to integrate Baselight M, choosing its prec...

25/11/2025

Inside the Secret World of Hollywood's Master Colourists

Article courtesy of The Hollywood Reporter Read the article Once hidden in post-production suites, the artists who make movies and TV shows look the way they ...

25/11/2025

FilmLight Colour Awards The Winners

Article courtesy of Deadline Read the article The Brutalist' & Bad Bunny's Nuevayol' Music Video Among 2025 FilmLight Colour Award Winners - Cam...

25/11/2025

FLUX.2 Image Generation Models Now Released, Optimized for NVIDIA RTX GPUs

Black Forest Labs - the frontier AI research lab developing visual generative AI models - today released the FLUX.2 family of state-of-the-art image generation ...

24/11/2025

HBO's The Shuffle' Reveals Longtime Connection of Sports and Entertainment

HBO's The Shuffle' Reveals Longtime Connection of Sports and Entertainm...

24/11/2025

2025 Sports Broadcasting Hall of Fame: Hiroshi Kiriyama, Sony Broadcast (and Industry) Technology Icon

2025 Sports Broadcasting Hall of Fame: Hiroshi Kiriyama, Sony Broadcast (and Ind...

24/11/2025

SVG Summit 2025 Preview: FIFA, NBC Olympics, Fox Sports, CBS Sports, Netflix, NFL, NBA, MLB, USTA Power Dec. 16 Conversations

SVG Summit 2025 Preview: FIFA, NBC Olympics, Fox Sports, CBS Sports, Netflix, NF...

24/11/2025

Case Study: YES Network Streamlines Broadcast Operations with Beam Dynamics

Case Study: YES Network Streamlines Broadcast Operations with Beam DynamicsBy SVG Staff Monday, November 24, 2025 - 11:18 am Print This Story | Subscribe ...

24/11/2025

Platinum White Paper: More of Everything: How Broadcasters are Changing Their Approach to Meet Rises in Consumer Demand with Calrec

Platinum White Paper: More of Everything: How Broadcasters are Changing Their Ap...

24/11/2025

Versant Media USA Sports President Matt Hong on How Versant Has Best of Both Worlds: a Start-Up Mentality and $7 Billion Revenues

Versant Media USA Sports President Matt Hong on How Versant Has Best of Both Wor...

24/11/2025

SVG Sit-Down: NABA Director-General Rebecca Hanson on How FCC's C-Band Auction Will Impact Broadcasters

SVG Sit-Down: NABA Director-General Rebecca Hanson on How FCC's C-Band Aucti...

24/11/2025

SVG New Sponsor Spotlight: Bolin Technology's Sapan Doshi on the Proliferation of PTZ Cameras for Sports Venues and Broadcasters

SVG New Sponsor Spotlight: Bolin Technology's Sapan Doshi on the Proliferati...

24/11/2025

Spotify and Acne Studios Welcome Robyn Back to the Stage in Los Angeles

Robyn made her long-awaited return to the stage this week, as Spotify and Acne Studios brought friends and top fans together for an unforgettable evening at the...

24/11/2025

Bara Is Back: The New Spotify Camp Nou Opens Its Gates

After more than two years of redevelopment, FC Barcelona returned to its spiritual home on November 22, hosting Athletic Club in the first La Liga match at the ...

24/11/2025

L3Harris' Next-Generation Weather Imager Ready to Deliver Life-Saving Weather Data Under Critical NOAA Satellite Program

The L3Harris next-generation imager for NOAA's GeoXO satellite system will c...

24/11/2025

JioStar and Nielsen Unveil Breakthrough Cross-Screen Measurement Study, Redefining Advertising Effectiveness in Live Sports

Mumbai - November 24, 2025 - In a first-of-its-kind initiative, JioStar, in coll...

24/11/2025

Fall Sports Plus Fresh Broadcast Slate Equals Big Gains, Reshuffled Company Rankings in Nielsen's October Media Distributor Gauge

Disney Achieves Largest Monthly Share Increase, Followed by FOX and Paramount, w...

24/11/2025

Sinclair Promotes Sean LaRose to VP and General Manager in Rochester

ROCHESTER, N.Y. Sinclair said it has elevated Sean LaRose, director of sales at WUHF and partner station WHAM here, to vice president and general manager, effec...

24/11/2025

Chaos and Connection: Meet the Unfiltered Cast of Reality Dating Series Badly in Love' as Trailer Debuts

Back to All News Chaos and Connection: Meet the Unfiltered Cast of Reality Dati...