Sony Pixel Power calrec Sony

How Deep Learning Is Aiding Preservation of Seneca and Other Endangered Languages

03/01/2019

Linguists estimate that at least half of the world's estimated 7,000 spoken languages will become extinct by the century's end, due to forces ranging from globalization to cultural assimilation.

Part of the challenge of documenting and revitalizing endangered languages is a lack of texts and speech recordings to work with. Seneca, a language of one of the six Iroquois Nations in North America, has only about 100 first-language speakers and several hundred more second-language learners.

Automatic speech recognition (ASR) technology is widely used to transcribe languages with millions or billions of speakers, like English and Mandarin. But it has only scratched the surface with languages like Seneca, which have vastly fewer speakers and significantly less data to work with.

Now a team of researchers at the Rochester Institute of Technology in New York, along with colleagues from the University at Buffalo, is tapping deep learning to bolster the ability of ASR. And while its focus is on Seneca, the project's vision encompasses the preservation of languages globally as well as an important part of our shared cultural history.

Knowing about different languages teaches us a lot about how our brain works, said Emily Prud'hommeaux, an assistant professor of computer science at Boston College and a research faculty member at RIT. When you document a language, you're preserving information not only about that language but also about how humans use language in general.

It's no coincidence that Prud'hommeaux and her team started with the Seneca language. Three members of the Seneca nation are part of the effort - a direct connection that is rare in research of this type, she said.

Leading the charge is Robbie Jimerson, a Ph.D. student in RIT's Golisano College of Computing and Information Science. He is a member of the Seneca Nation of Indians and is passionate about ensuring the survival of the Seneca language.

There's a big effort by the leaders of the tribe to preserve and promote our language, said Jimerson. I was looking for an opportunity to contribute.

Using GANs to Create More Language Samples Now in its third year, the project has had challenges when it comes to accumulating language data. Jimerson said the Seneca community can be guarded about what it shares with other people, so there wasn't an abundance of recordings of the language being spoken. He set out to change that.

He started by recording friends and elders who speak the language and asking them to record their friends. He found out whenever someone was speaking Seneca in public. He asked for family recordings of elders telling stories handed down from previous generations. And he grabbed any publicly available videos or recordings he could find online.

The team has fine-tuned an ASR model for Seneca, running it through generative adversarial networks to create more samples out of the limited number of recordings. The model turns wave files of the spoken language into streams of characters, while computing probability and making corrections.

The resulting data is fed into a deep learning model that in turn expands upon the ASR model's accuracy.

The team's networks run in two compute settings: on a nine-server machine learning lab running a variety of NVIDIA Tesla GPUs, and on a university cluster of large servers, each running 10 NVIDIA Tesla P4 GPUs. Each cluster runs a range of deep learning frameworks such as TensorFlow and Caffe.

The computer engineering cluster is for all students in the computer engineering department, and so they have to compete' for these resources, said Ray Ptucha, assistant professor of computer engineering at RIT, another collaborator on this project.

With access to these clusters at a premium, Jimerson tests code and checks the stability of models on a local machine running an NVIDIA TITAN X rather than inconvenience other students by running a model that might crash.

Achieving Better Accuracy So far, the team's efforts have brought the word error rate of its ASR model from 70 percent down to 56 percent. The goal, said Prud'hommeaux, is to get that rate down to 25 percent, which is where ASR systems were in processing English several years ago.

The more samples of spoken and written Seneca the team can accumulate, the more the error rate will decrease. (Today, English ASR models can achieve word error rates as low as 5 percent.)

The team's work is expected to help with language preservation efforts around the world.

Prud'hommeaux said the team has an agreement with an archiving institution that's a condition of a grant the project received from the National Science Foundation. The resulting language archiving database will be made available as a resource for other efforts seeking to document threatened languages.

Additionally, Prud'hommeaux said the team's work could prove helpful for any deep learning effort that has to make do with limited amounts of data.

Read more about the team's work in their research papers here and here.

Feature image: The Haudenosaunee (Iroquois Confederacy) flag, via Wikimedia Commons.
LINK: https://blogs.nvidia.com/blog/2019/01/02/deep-learning-preserves-senec...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

03/04/2026

2026 NAB Show Exhibitor Insight: Techex

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

TelevisaUnivision Signs New Nielsen Media Intelligence Deal

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

The WNET Group, JIB Launch NHK World-Japan in New York

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

NAB Leadership Foundation Welcomes New Board Members

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

EverPass Media Expands Distribution Deal with Netflix

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

Versant Acquires AI-Data Platform StockStory

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

CVP Grows European Footprint with Strategic Expansion in...

CVP, one of Europe's leading suppliers of professional video and broadcast solutions, today announces the launch of its new German operation and the formati...

03/04/2026

MRMC Announces Appointment of Chief Operating Officer

Mark Roberts Motion Control (MRMC) today announces the appointment of Nick Barthee as Chief Operating Officer, strengthening its leadership as the company conti...

03/04/2026

Net Insight Introduces Programmable Trust Boundaries for...

Net Insight introduces programmable Trust Boundaries that make live media interconnection predictable as traffic moves between facilities, networks and cloud en...

03/04/2026

Winning in the new media economy: Avid showcases AI-powered, connected intelligence to unlock media value at NAB Show 2026

Winning in the new media economy: Avid showcases AI-powered, connected intellige...

03/04/2026

NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibility and Loudness at NAB 2026

NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibilit...

03/04/2026

NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances

NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances Brie Clayton April 2, 2026 0 Comments PlayBox Neo will showcase the latest i...

03/04/2026

For Taku Hirano, Everything Is Connected

For Taku Hirano, Everything Is Connected From touring and composition to teaching and instrument design, the in-demand percussionist sees it all as one body o...

03/04/2026

Berklee Honors Humberto Ramirez with Master of Latin Music Award

Berklee Honors Humberto Ramirez with Master of Latin Music Award The alumnus and acclaimed trumpeter is honored for his influence as a performer, composer, an...

02/04/2026

HBO and NFL Films Announce Hard Knocks: Training Camp with the Seattle Seahawks, Debuting August 11

HBO and NFL Films have announced Hard Knocks: Training Camp with the Seattle Sea...

02/04/2026

NAB 2026: Haivision Unveils Makito ONE Video Transport Platform

Haivision has announced the Makito ONE, a single-blade video encoding and decoding platform, at NAB Show 2026. The platform combines dual-channel video encoding...

02/04/2026

NAB 2026: Telestream Introduces UP.Lens Cloud-Based Multiviewer and Monitoring Service

Telestream has introduced UP.Lens, a cloud-based multiviewer and monitoring serv...

02/04/2026

NAB 2026: MRMC to Showcase Robotic Camera Technology and Mark 60th Anniversary

Mark Roberts Motion Control (MRMC) will exhibit at NAB Show 2026 (Booth C5220, April 19-22, Las Vegas Convention Center), marking the company's 60th anniver...

02/04/2026

NAB 2026: Net Insight Introduces Programmable Trust Boundaries for Live Media Interconnection

Net Insight has introduced programmable Trust Boundaries, a feature integrated i...

02/04/2026

NAB 2026: Bitmovin Adds SGAI Support to Playback Products

Bitmovin has announced support for SGAI (Server-Guided Ad Insertion) in its playback products, using HLS interstitials. SGAI combines elements of client-side an...

02/04/2026

Binghamton University Athletics Adds Riedel SimplyLive RiMotion R12 for Student-Run Productions

Riedel Communications' SimplyLive RiMotion R12 replay system is supporting B...

02/04/2026

NAB 2026: LTN and Ateme Announce Integration of Video Processing with IP Transport

LTN, a managed IP video transport company, and Ateme, a video compression and de...

02/04/2026

TDF Expands Channel Capacity on Terrestrial Broadcast Network with Harmonic

Harmonic has announced that TDF, a broadcast infrastructure operator in France, has deployed Harmonic's XOS Advanced Media Processor and ProStream X Video S...

02/04/2026

United Rugby Championship Reports First-Year Results with Eluvio Streaming Platform

Eluvio and the United Rugby Championship (URC) have announced first-year results...

02/04/2026

Kansas City Current and Scripps Sports Announce ION as Broadcast Home of 2026 Teal Rising Cup

The Kansas City Current and Scripps Sports have announced that ION will broadcas...

02/04/2026

ESPN Announces Courtside Alt-Cast for Women's Final Four

ESPN will debut Courtside at the Women's Final Four Presented by AT&T, an alt-cast airing Friday, April 3 at 7 p.m. and 9:30 p.m. ET on ESPN2, and Sunday, A...

02/04/2026

PAMA and Shure Accept Applications for 6th Annual Mark Brunner Professional Audio Scholarship

The Professional Audio Manufacturers Alliance (PAMA) and Shure Incorporated are ...

02/04/2026

ESPN's MegaCast Coverage of 2026 NCAA Women's Final Four Begins Friday, April 3 in Phoenix

ESPN's MegaCast Coverage of 2026 NCAA Women's Final Four Begins Friday, ...

02/04/2026

DAZN Launches Playmakers Creator Program

DAZN has announced the launch of DAZN Playmakers, a global influencer program designed to build a network of sports content creators. The programme will give cr...

02/04/2026

NFL Network Unveils New Production Ops Leadership Structure as ESPN Takes Over

Tony Cole, Jessica Lee shift into new roles reporting to ESPN SVP/Content Operations Chris Calcinari....

02/04/2026

TNT Sports and CBS Sports to Reunite Michigan's Iconic Fab Five for Special NCAA Men's Final Four Altcast on truTV & HBO Max

Michigan's Fab Five will reunite for an alternate presentation of the Mich...

02/04/2026

Coming of Age: ESPN and NHL's Inside Out Classic Marks Another Step Forward for the Animated Alternative Broadcast

Real-time tracking, virtual production, and Pixar storytelling converge for Apri...

02/04/2026

Streaming Around the Moon: NASA+ Goes Live From Historic Artemis II Mission

NASA's long-awaited Artemis II mission has launched four astronauts on a 10-day journey around the moon, marking the first manned launch toward the moon sin...

02/04/2026

Release Rundown: What to Watch in April, From Bunnylovr to Omaha

(L-R) Molly Belle Wright, Wyatt Solis, and John Magaro appear in Omaha by Cole Webley, an official selection of the 2025 Sundance Film Festival. (Photo courte...

02/04/2026

VEMIA Auction incoming

4 - 11 April 2026 VEMIA's latest gear auction is just around the corner, and there's already a wealth of sought-after instruments and studio gear li...

02/04/2026

Universal Audio's Voice Of God now native

Little Labs emulation now available to all The latest Universal Audio plug-in to become available outside of the company's DSP-powered UAD2 platform has...

02/04/2026

SBS Welcomes Back Cup Fever! For The Biggest FIFA World Cup Ever

SBS Welcomes Back Cup Fever! For The Biggest FIFA World Cup Ever 2 April, 2026 Media releases Santo Cilauro, Ed Kavalee and a host of special guests retur...

02/04/2026

Truth. Power. Perspective: NITV's Flagship Current Affairs Programs Return to Lead the National Conversation

Truth. Power. Perspective: NITV's Flagship Current Affairs Programs Return t...

02/04/2026

L3Harris Powers First Crewed Mission Around the Moon in 50 Years

L3Harris has successfully powered the historic launch of the Artemis II mission, providing propulsion and avionics....

02/04/2026

Scripps Completes Sale of WRTV to Circle City Broadcasting

Share Copy link Facebook X Linkedin Bluesky Email...

02/04/2026

GoVertical! AiDi Powers Real-Time 9:16 Autocropping for I...

Already deployed extensively by NBC Sports, FOR-A Corporation will demonstrate GoVertical! AiDi, the real-time 9:16 autocropping feature of viztrick AiDi, durin...

02/04/2026

Elite Media Technologies Selects Interra Systems BATON Fi...

Interra Systems, a provider of end-to-end quality assurance solutions for the digital media industry, announced that Elite Media Technologies has selected its B...

02/04/2026

TDF Expands Broadcast Channel Lineup with Harmonic

Harmonic's Media Processing Solutions Maximize Bandwidth Efficiency for Terrestrial Broadcast Delivery Harmonic (NASDAQ: HLIT) today announced that TDF, a...

02/04/2026

FOR-A's Software-Defined, AI-Powered Development Advances...

NBC Sports Deploys viztrick AiDi to Stream Live Events in 9:16 Mobile-First Formats with Auto Tracking, Development Signals Strategic Shift for FOR-A Long reco...

02/04/2026

Evergent showcases innovations in sports streaming and mo...

Evergent will showcase new innovations in subscriber lifecycle management and monetization at NAB Show 2026 (Las Vegas, April 18 22), including: New advances i...

02/04/2026

Binghamton University Strengthens Student Run Productions...

Riedel Communications is proud to be part of Binghamton University, State University of New York, Athletics' milestone year, celebrating the university'...