Sony Pixel Power calrec Sony

Now Hear This: World's Most Flexible Sound Machine Debuts

25/11/2024

A team of generative AI researchers created a Swiss Army knife for sound, one that allows users to control the audio output simply using text.

While some AI models can compose a song or modify a voice, none have the dexterity of the new offering.

Called Fugatto (short for Foundational Generative Audio Transformer Opus 1), it generates or transforms any mix of music, voices and sounds described with prompts using any combination of text and audio files.

For example, it can create a music snippet based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice - even let people produce sounds never heard before.

This thing is wild, said Ido Zmishlany, a multi-platinum producer and songwriter - and cofounder of One Take Audio, a member of the NVIDIA Inception program for cutting-edge startups. Sound is my inspiration. It's what moves me to create music. The idea that I can create entirely new sounds on the fly in the studio is incredible.

A Sound Grasp of Audio We wanted to create a model that understands and generates sound like humans do, said Rafael Valle, a manager of applied audio research at NVIDIA and one of the dozen-plus people behind Fugatto, as well as an orchestral conductor and composer.

Supporting numerous audio generation and transformation tasks, Fugatto is the first foundational generative AI model that showcases emergent properties - capabilities that arise from the interaction of its various trained abilities - and the ability to combine free-form instructions.

Fugatto is our first step toward a future where unsupervised multitask learning in audio synthesis and transformation emerges from data and model scale, Valle said.

A Sample Playlist of Use Cases For example, music producers could use Fugatto to quickly prototype or edit an idea for a song, trying out different styles, voices and instruments. They could also add effects and enhance the overall audio quality of an existing track.

The history of music is also a history of technology. The electric guitar gave the world rock and roll. When the sampler showed up, hip-hop was born, said Zmishlany. With AI, we're writing the next chapter of music. We have a new instrument, a new tool for making music - and that's super exciting.

An ad agency could apply Fugatto to quickly target an existing campaign for multiple regions or situations, applying different accents and emotions to voiceovers.

Language learning tools could be personalized to use any voice a speaker chooses. Imagine an online course spoken in the voice of any family member or friend.

Video game developers could use the model to modify prerecorded assets in their title to fit the changing action as users play the game. Or, they could create new assets on the fly from text instructions and optional audio inputs.

Making a Joyful Noise One of the model's capabilities we're especially proud of is what we call the avocado chair, said Valle, referring to a novel visual created by a generative AI model for imaging.

For instance, Fugatto can make a trumpet bark or a saxophone meow. Whatever users can describe, the model can create.

With fine-tuning and small amounts of singing data, researchers found it could handle tasks it was not pretrained on, like generating a high-quality singing voice from a text prompt.

Users Get Artistic Controls Several capabilities add to Fugatto's novelty.

During inference, the model uses a technique called ComposableART to combine instructions that were only seen separately during training. For example, a combination of prompts could ask for text spoken with a sad feeling in a French accent.

The model's ability to interpolate between instructions gives users fine-grained control over text instructions, in this case the heaviness of the accent or the degree of sorrow.

I wanted to let users combine attributes in a subjective or artistic way, selecting how much emphasis they put on each one, said Rohan Badlani, an AI researcher who designed these aspects of the model.

In my tests, the results were often surprising and made me feel a little bit like an artist, even though I'm a computer scientist, said Badlani, who holds a master's degree in computer science with a focus on AI from Stanford.

The model also generates sounds that change over time, a feature he calls temporal interpolation. It can, for instance, create the sounds of a rainstorm moving through an area with crescendos of thunder that slowly fade into the distance. It also gives users fine-grained control over how the soundscape evolves.

Plus, unlike most models, which can only recreate the training data they've been exposed to, Fugatto allows users to create soundscapes it's never seen before, such as a thunderstorm easing into a dawn with the sound of birds singing.

A Look Under the Hood Fugatto is a foundational generative transformer model that builds on the team's prior work in areas such as speech modeling, audio vocoding and audio understanding.

The full version uses 2.5 billion parameters and was trained on a bank of NVIDIA DGX systems packing 32 NVIDIA H100 Tensor Core GPUs.

Fugatto was made by a diverse group of people from around the world, including India, Brazil, China, Jordan and South Korea. Their collaboration made Fugatto's multi-accent and multilingual capabilities stronger.

One of the hardest parts of the effort was generating a blended dataset that contains millions of audio samples used for training. The team employed a multifaceted strategy to generate data and instructions that considerably expanded the range of tasks the model could perform, while achieving more accurate performance and enabling new tasks without requiring additional data.

They also scrutinized existing datasets to reveal new relationships among the dat
LINK: https://blogs.nvidia.com/blog/fugatto-gen-ai-sound-model/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

10/04/2026

The Invisible OPEX Killer: Is Your Server Room Dragging You Down?

The Invisible OPEX Killer: Is Your Server Room Dragging You Down? In the broadcast world, we talk a lot about uptime. We talk about talent retention, latency...

10/04/2026

NAB 2026: Imagine Communications to Showcase Expanded Multiviewer Portfolio

Imagine Communications will showcase its multiviewer portfolio at NAB Show 2026 (April 19-22, Booth N1328, Las Vegas Convention Center), including Prismon and t...

10/04/2026

NAB 2026: Chyron Releases PRIME VSAR 2.3 with Updated Unreal Engine Integration

Chyron has released PRIME VSAR 2.3, an update to its virtual set and augmented reality solution for broadcast. The release adds compatibility with Unreal Engine...

10/04/2026

NAB 2026: Techex to Showcase New tx darwin Capabilities

Techex will exhibit at NAB Show 2026 (Booth W2267, April 19-23, Las Vegas Convention Center), demonstrating new tx darwin features including consumer multiview,...

10/04/2026

NAB 2026: NDI to Showcase Ecosystem and NDI 6.3

NDI will exhibit at NAB Show 2026, demonstrating its IP video ecosystem through live partner integrations, NDI 6.3 features, AI metadata workflows, and creator ...

10/04/2026

FOR-A Acquires Tamura Corporations Information Equipment Business

FOR-A has announced the acquisition of all shares of Tamu Radiance Corporation, a new company spun off from the Information Equipment Business of Tamura Corpora...

10/04/2026

NAB 2026: InSync Technology to Unveil New Video Processing and Frame Rate Conversion Products

InSync Technology will showcase new and updated video conversion products at NAB...

10/04/2026

TNT Sports and DAZN Announce Monthly Boxing Event Series in the United States

TNT Sports and DAZN have announced a partnership to air monthly boxing events in the United States under the brand The Fight. The series will be promoted in p...

10/04/2026

Panasonic Introduces SQ3 Series 4K LCD Displays for Professional Environments

Panasonic Projector and Display has announced the SQ3 Series of 4K LCD displays as part of its MEVIX professional display portfolio. All sizes will be available...

10/04/2026

Amagi Adds Agentic Capabilities to Its Media Operations Platform

Amagi has announced the addition of Agentic Media Operations to its Amagi NOW platform, integrating AI reasoning agents across its media supply chain workflows ...

10/04/2026

LTN Announces Network Enhancements Ahead of C-Band Spectrum Auction

LTN has announced enhancements to its global IP video network targeting broadcasters transitioning from satellite distribution. The updates come ahead of US fed...

10/04/2026

Daktronics Installs New LED Displays at Yankee Stadium

Daktronics has installed new LED displays at Yankee Stadium, upgrading the main centerfield board, two flanking boards, and two ribbon displays spanning the 200...

10/04/2026

NAB 2026: Harmonic Announces AI and Cloud Updates to Hybrid Streaming Solution

Harmonic has announced updates to its hybrid streaming solution, including Model Context Protocol (MCP) connectivity for AI applications, cloud-native deploymen...

10/04/2026

NAB 2026: MultiDyne to Debut FiberSaver-10G and VF-9100

MultiDyne Video and Fiber Optic Systems will introduce two new fiber transport products at NAB Show 2026 (Booth C4425, April 19-22): the FiberSaver-10G waveleng...

10/04/2026

NAB 2026: Telos Alliance and ip-studio to Demonstrate STUDIO ZERO

Telos Alliance and ip-studio will demonstrate STUDIO ZERO, a cloud-hosted virtual studio, at NAB Show 2026. First introduced at NAB Show 2023, STUDIO ZERO integ...

10/04/2026

Pixotope and d&b Solutions Announce Strategic Partnership for XR and Virtual Studio Production

d&b solutions, a London-based audio-visual, lighting, and media integration grou...

10/04/2026

ARRI and SmallHD Announce Lens Data Monitor Overlay License for Hi-5 and Hi-5 SX

ARRI and SmallHD have announced a new expansion license for ARRI's Hi-5 and Hi-5 SX hand units that displays lens data overlays on supported SmallHD monitor...

10/04/2026

Roku to Stream Exclusive Savannah Bananas Game Package on Roku Sports Channel

Roku and the Banana Ball Championship League (BBCL) have announced an exclusive streaming partnership to bring five BBCL games to the Roku Sports Channel in 202...

10/04/2026

Ratings Roundup: More Than 18 Million Fans Tune Into 2026 NCAA Mens March Madness on TNT and CBS Sports

Ratings Roundup is a rundown of recent rating news and is derived from press rel...

10/04/2026

No Other Land, Mr. Nobody Against Putin,and More Sundance Institute-Supported Films Nominated for Peabody Awards

The Peabody Awards don't just recognize great storytelling, they spotlight t...

10/04/2026

Fans Crown Winners at the First Spotify Podcast Awards in France

After launching the Spotify Podcast Awards in Mexico last year, we brought the fan-voted celebration to Paris this week for its first edition in France. Hosted ...

10/04/2026

Yamaha launch the DXR/DXS & CXR/CXS Mk3

Powered and unpowered live PA ranges upgraded Yamaha have just refreshed four of their hugely popular PA speaker ranges, delivering significant improvements...

10/04/2026

UJAM open Gorilla Engine to third-party developers

Underlying plug-in & VI technology now available to others UJAM's latest announcement sees the company open up' Gorilla Engine, the development pla...

10/04/2026

2026 NAB Show Exhibitor Insight: Bitcentral

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Bitcentral To Feature Connected Media Workflows At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Bitcentral to Showcase Connected Media Workflows and Inte...

NEWPORT BEACH, Calif., April 10, 2026 Bitcentral, a leading provider of professional media solutions for broadcast and digital video, will showcase its latest...

10/04/2026

Ikegami to Introduce Expanded Range of Broadcast Production Solutions at NAB 2026

Ikegami to Introduce Expanded Range of Broadcast Production Solutions at NAB 202...

10/04/2026

AJA Debuts SMPTE ST 2110 and openGear Solutions Ahead of NAB 2026

AJA Debuts SMPTE ST 2110 and openGear Solutions Ahead of NAB 2026 Brie Clayton April 10, 2026 0 Comments New gear and updates address evolving hybrid ...

10/04/2026

Portland Fire+ Streaming Platform Launches

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Tod Musgrave Joins Proton as U.S. Sales & Marketing Director

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Proton Expands Minicam Portfolio With Proton Pro At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

FCC To Vote on Changes to Audible Crawl Rule

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Frequency Launches AI Platform for Streaming Television a...

Frequency, the engine behind the worlds leading streaming television channels, today launched its AI platform for Frequency Studio, powering the entire channel ...

10/04/2026

UKTV Highlights: Saturday April 25th - Friday May 1st 2026

What can I watch on UKTV and stream on U this week? This week on UKTV and the free streaming service U, viewers can watch a range of new and returning programm...

10/04/2026

Jnger Audio Joins EBU ADM Integration Group as Founding Member to Help Advance ADM/S-ADM Integration

J nger Audio Joins EBU ADM Integration Group as Founding Member to Help Advance...

10/04/2026

Multi-BAFTA-winning Chernobyl makes its free-to-air debut, marking 40 years since the disaster

Five-part Sky Original drama airs nightly on Sky Mix and Sky Atlantic from 20 Ap...

09/04/2026

Yospace surpasses 10 billion ads stitched in a single month, as ad-supported streaming surges

Staines-upon-Thames, UK, 09, April, 2026 - Yospace, the trusted leader in Dynam...

09/04/2026

just:play pro 2026 and just:live pro 2026 Sneak Preview News for NAB 2026

just:play pro 2026 and just:live pro 2026 Sneak Preview News for NAB 2026 More Details:At NAB 2026, ToolsOnAir will showcase just:play pro 2026 and just:live p...

09/04/2026

just:in mac pro 2026 - The Next Level of Professional Recording on macOS at NAB 2026

just:in mac pro 2026 - The Next Level of Professional Recording on macOS at NAB ...

09/04/2026

NAB 2026: Zixi to Demonstrate Live Video Workflows and Satellite Replacement

Zixi will demonstrate IP-based live video workflow solutions at NAB Show 2026 (Booth W2057). The industry is moving quickly toward IP-based distribution as br...

09/04/2026

Deloitte Research: Women's Elite Sports Revenues Expected to Reach at Least $3 Billion in 2026

Global women's elite sports revenues are expected to reach at least $3 billi...

09/04/2026

Monitor Engineer Gavin Tempany Mixes Kylie Minogue's Tension Tour on Solid State Logic L550 Plus

Monitor engineer Gavin Tempany mixed Kylie Minogue s Tension Tour on a Solid Sta...

09/04/2026

NAB 2026: KOKUSAI DENKI Electric America to Debut New 4K Camera and Remote Control Panel

KOKUSAI DENKI Electric America will exhibit at NAB Show 2026 (Booth C5507), debu...

09/04/2026

NBC Sports Reviews Innovations and Milestones from Its 2025-26 NBA Regular Season

With the 2025-26 NBA regular season concluded and the playoffs beginning next we...