
University of Washington researchers have developed new algorithms that can turn audio clips into a realistic, lip-synced video of the person speaking those words.
As detailed in a paper to be presented August 2 at SIGGRAPH 2017 in L.A., the team successfully generated realistic video of former president Barack Obama talking about terrorism, fatherhood, job creation and other topics using audio clips of those speeches and existing weekly video addresses that were originally on a different topic.
Ira Kemelmacher-Shlizerman, an assistant professor at the UW's Paul G. Allen School of Computer Science & Engineering said, Realistic audio-to-video conversion has practical applications like improving video conferencing for meetings, as well as futuristic ones such as being able to hold a conversation with a historical figure in virtual reality by creating visuals just from audio.
In a visual form of lip-syncing, the system converts audio files of an individual's speech into realistic mouth shapes, which are then grafted onto and blended with the head of that person from another existing video.
In the future video, chat tools like Skype or Messenger will enable anyone to collect videos that could be used to train computer models, Kemelmacher-Shlizerman said.
Because streaming audio over the internet takes up far less bandwidth than video, the new system has the potential to end video chats that are constantly timing out from poor connections.
When you watch Skype or Google Hangouts, often the connection is stuttery and low-resolution and really unpleasant, but often the audio is pretty good, said co-author and Allen School professor Steve Seitz. So if you could use the audio to produce much higher-quality video, that would be terrific.
By reversing the process feeding video into the network instead of just audio the team could also potentially develop algorithms that could detect whether a video is real or manufactured.
The new machine learning tool makes significant progress in overcoming what's known as the uncanny valley problem, which has dogged efforts to create realistic video from audio. When synthesised human likenesses appear to be almost real but still manage to somehow miss the mark people find them creepy or off-putting.
People are particularly sensitive to any areas of your mouth that don't look realistic, said lead author Supasorn Suwajanakorn, a recent doctoral graduate in the Allen School. If you don't render teeth right or the chin moves at the wrong time, people can spot it right away and it's going to look fake. So you have to render the mouth region perfectly to get beyond the uncanny valley.
A neural network first converts the sounds from an audio file into basic mouth shapes. Then the system grafts and blends those mouth shapes onto an existing target video and adjusts the timing to create a new realistic, lip-synced video.
Previously, audio-to-video conversion processes have involved filming multiple people in a studio saying the same sentences over and over to try to capture how a particular sound correlates to different mouth shapes, which is expensive, tedious and time-consuming. By contrast, Suwajanakorn developed algorithms that can learn from videos that exist in the wild on the internet or elsewhere.
There are millions of hours of video that already exist from interviews, video chats, movies, television programs and other sources. And these deep learning algorithms are very data hungry, so it's a good match to do it this way, Suwajanakorn said.
Rather than synthesising the final video directly from audio, the team tackled the problem in two steps. The first involved training a neural network to watch videos of an individual and translate different audio sounds into basic mouth shapes.
By combining previous research from the UW Graphics and Image Laboratory team with a new mouth synthesis technique, they were then able to realistically superimpose and blend those mouth shapes and textures on an existing reference video of that person. Another key insight was to allow a small time shift to enable the neural network to anticipate what the speaker is going to say next.
The new lip-syncing process enabled the researchers to create realistic videos of Obama speaking in the White House, using words he spoke on a television talk show or during an interview decades ago.
Currently, the neural network is designed to learn on one individual at a time, meaning that Obama's voice speaking words he actually uttered is the only information used to drive the synthesised video. Future steps, however, include helping the algorithms generalise across situations to recognise a person's voice and speech patterns with less data with only an hour of video to learn from, for instance, instead of 14 hours.
The research was funded by Samsung, Google, Facebook, Intel and the UW Animation Research Labs.
A neural network first converts the sounds from an audio file into basic mouth shapes. Then the system grafts and blends those mouth shapes onto an existing target video and adjusts the timing to create a new realistic, lip-synced video.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
01/04/2026
January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION
Douyin Users Can Now Create And Share Videos With Stun...
21/02/2026
With Software Defined Broadcasting more established in Milan Cortina look for Los Angeles 2028 to have less hardware and more cloud-based software systems...
21/02/2026
The SVP of Olympic Operations on turning CAD drawings into reality, building tru...
21/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/02/2026
Back to All News
Netflix Unveils the Trailer of Accused', A Psychological ...
20/02/2026
Gravity Media and Los Angeles-based Green Couch Entertainment announce a strateg...
20/02/2026
IMAX announces it is working with Apple TV to bring the 2026 FIA Formula One Wor...
20/02/2026
Daktronics has partnered with the Philadelphia Phillies to design, manufacture, ...
20/02/2026
ESPN announces the upcoming launch of Women's Sports Sundays - a first-of-it...
20/02/2026
As the Seattle Seahawks and New England Patriots faced off in the NFL's biggest sporting event of the season on Sun., Feb. 8, Sennheiser wireless solutions ...
20/02/2026
ESPN announces its 2026 Major League Baseball spring training schedule, which includes four national games on ESPN, six games on ESPN Unlimited, and more than 2...
20/02/2026
Open Broadcast Systems, which specializes in software-based professional video transport, has added support for 200 Gigabit Ethernet to its range of encoders an...
20/02/2026
Chyron announces the release of PAINT 10.3, which is designed to help analysts and operators turn live action into clearer, faster on-air storytelling.
PAINT 1...
20/02/2026
With full squad workouts underway, MLB Network's live Spring Training game s...
20/02/2026
Tech enhancements, marquee productions are expected to take advantage of a summe...
20/02/2026
In-venue and creative video staffers at the professional and collegiate level ha...
20/02/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
20/02/2026
Speaking with SVG Europe after one of Team GB's greatest days at a Winter Olympics, BBC Sport's head of major events, Ron Chakraborty, explains the broa...
20/02/2026
Making Winter Games Olympic magic is the goal for every broadcaster in Italy cov...
20/02/2026
Curling, one of the least-dangerous Winter Olympic sports, is dominating the Mil...
20/02/2026
BBC Sport's presence at the 2026 Winter Games is centred around a significan...
20/02/2026
BBC Sport is bringing together its linear TV and streaming digital arms in a str...
20/02/2026
To broaden the appeal of winter sports at Milano Cortina, the BBC has integrated...
20/02/2026
Just in time for the start of Apple TV's inaugural season as the exclusive U...
20/02/2026
One big challenge was to depict the character of each of very different and wide...
20/02/2026
(L-R) Writer-director Amanda Kramer photographs the photographers at the premiere of her film By Design at the Library Center Theatre in Park City. (Photo by ...
20/02/2026
In our latest blog, Tim Pearson explores the impact that increased memory prices are having on the consumer electronics market, and particularly the set-top box...
20/02/2026
Calrec Type R: Shaping the Future of Radio from the Heart of Flirt FM
Love may have filled the airwaves last week for Valentine's Day, and we've just c...
20/02/2026
NEW YORK - February 10, 2026 - An estimated 125.6* million viewers watched Super Bowl LX on Sunday, February 8, according to Nielsen's Big Data Panel meas...
20/02/2026
NEW YORK - February 19, 2026 - Nielsen today shared updated and final Super Bowl...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/02/2026
A leading global investment bank, with offices at Two International Finance Centre in Hong Kong, partnered with systems integrators Global Vision Engineering (G...
20/02/2026
Rise AV and Rise Broadcast, the global not-for-profit organisations dedicated to improving gender diversity across technical industries, have today announced a ...
20/02/2026
Open Broadcast Systems, the leader in software-based professional video transport, has added support for 200 Gigabit Ethernet to its range of encoders and decod...
20/02/2026
Signiant today announced the formation of its Customer Advisory Board (CAB), bringing together a select group of customers to collaborate on product strategy, r...
20/02/2026
PTZOptics today announced the launch of its Visual Reasoning initiative that makes video more actionable by combining robotic PTZ camera systems, AI, and open i...
20/02/2026
Amino, a global media technology provider delivering devices, software and cloud services that simplify and elevate video delivery, today announced the successf...
20/02/2026
SMPTE , the home of media professionals, technologists, and engineers, today announced its call for technical papers for the SMPTE 2026 Media Technology Summit....
20/02/2026
Wowza Media Systems today announced that Granicus, a leading provider of digital engagement solutions for governments, continues to rely on Wowza to power its h...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...