
Linguists estimate that at least half of the world's estimated 7,000 spoken languages will become extinct by the century's end, due to forces ranging from globalization to cultural assimilation.
Part of the challenge of documenting and revitalizing endangered languages is a lack of texts and speech recordings to work with. Seneca, a language of one of the six Iroquois Nations in North America, has only about 100 first-language speakers and several hundred more second-language learners.
Automatic speech recognition (ASR) technology is widely used to transcribe languages with millions or billions of speakers, like English and Mandarin. But it has only scratched the surface with languages like Seneca, which have vastly fewer speakers and significantly less data to work with.
Now a team of researchers at the Rochester Institute of Technology in New York, along with colleagues from the University at Buffalo, is tapping deep learning to bolster the ability of ASR. And while its focus is on Seneca, the project's vision encompasses the preservation of languages globally as well as an important part of our shared cultural history.
Knowing about different languages teaches us a lot about how our brain works, said Emily Prud'hommeaux, an assistant professor of computer science at Boston College and a research faculty member at RIT. When you document a language, you're preserving information not only about that language but also about how humans use language in general.
It's no coincidence that Prud'hommeaux and her team started with the Seneca language. Three members of the Seneca nation are part of the effort - a direct connection that is rare in research of this type, she said.
Leading the charge is Robbie Jimerson, a Ph.D. student in RIT's Golisano College of Computing and Information Science. He is a member of the Seneca Nation of Indians and is passionate about ensuring the survival of the Seneca language.
There's a big effort by the leaders of the tribe to preserve and promote our language, said Jimerson. I was looking for an opportunity to contribute.
Using GANs to Create More Language Samples Now in its third year, the project has had challenges when it comes to accumulating language data. Jimerson said the Seneca community can be guarded about what it shares with other people, so there wasn't an abundance of recordings of the language being spoken. He set out to change that.
He started by recording friends and elders who speak the language and asking them to record their friends. He found out whenever someone was speaking Seneca in public. He asked for family recordings of elders telling stories handed down from previous generations. And he grabbed any publicly available videos or recordings he could find online.
The team has fine-tuned an ASR model for Seneca, running it through generative adversarial networks to create more samples out of the limited number of recordings. The model turns wave files of the spoken language into streams of characters, while computing probability and making corrections.
The resulting data is fed into a deep learning model that in turn expands upon the ASR model's accuracy.
The team's networks run in two compute settings: on a nine-server machine learning lab running a variety of NVIDIA Tesla GPUs, and on a university cluster of large servers, each running 10 NVIDIA Tesla P4 GPUs. Each cluster runs a range of deep learning frameworks such as TensorFlow and Caffe.
The computer engineering cluster is for all students in the computer engineering department, and so they have to compete' for these resources, said Ray Ptucha, assistant professor of computer engineering at RIT, another collaborator on this project.
With access to these clusters at a premium, Jimerson tests code and checks the stability of models on a local machine running an NVIDIA TITAN X rather than inconvenience other students by running a model that might crash.
Achieving Better Accuracy So far, the team's efforts have brought the word error rate of its ASR model from 70 percent down to 56 percent. The goal, said Prud'hommeaux, is to get that rate down to 25 percent, which is where ASR systems were in processing English several years ago.
The more samples of spoken and written Seneca the team can accumulate, the more the error rate will decrease. (Today, English ASR models can achieve word error rates as low as 5 percent.)
The team's work is expected to help with language preservation efforts around the world.
Prud'hommeaux said the team has an agreement with an archiving institution that's a condition of a grant the project received from the National Science Foundation. The resulting language archiving database will be made available as a resource for other efforts seeking to document threatened languages.
Additionally, Prud'hommeaux said the team's work could prove helpful for any deep learning effort that has to make do with limited amounts of data.
Read more about the team's work in their research papers here and here.
Feature image: The Haudenosaunee (Iroquois Confederacy) flag, via Wikimedia Commons.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
09/10/2026
September 10 2026, 06:00 (PDT) Dolby Expands Dolby OptiView Platform with New Capabilities at IBC 2026
New Sports Intelligence helps providers better unders...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
18/09/2026
DHD reports a high level of interest in its range of audio production equipment on all four days of the September 11-14 International Broadcasting Convention in...
18/09/2026
Sonnet Announces Thunderbolt 5 eGPU With 850-Watt Power Supply and Integrated Do...
18/09/2026
DeckBridge adds live Resolve playback and Fairlight controls to Stream Deck
Brie Clayton September 17, 2026
0 Comments
DeckBridge starting profiles ac...
18/09/2026
Learn to arrange tracks using Color With Live's Free Stencil Pack
Brie Clayton September 17, 2026
0 Comments
Stencils are Color With Live's or...
18/09/2026
Berklee Establishes Susan Tedeschi Scholarship in Honor of Grammy-Winning Alumna The scholarship is announced as Tedeschi returns to headline the Berklee Bloc...
18/09/2026
The inside scoop on attention-grabbing market assets from those in the know 17 September 2026
(L-R): Sophie Green, Todd Brown and Edwina Waddy
Find this episo...
17/09/2026
Lions vs. Bills will also be the first regular-season game at new Highmark Stadi...
17/09/2026
Teradek has announced Bolt 7, a wireless video and data system that adds an integrated 900MHz radio to the Bolt's zero-delay video pipeline. The 900MHz chan...
17/09/2026
DAZN has announced three appointments to its executive leadership team, all reporting to CEO Shay Segev effective immediately.
Patrick Delany has been appointe...
17/09/2026
Monumental Sports and Entertainment (MSE) has announced a leadership restructuring. Zach Leonsis has been named President and Vice Chairman, a newly created pos...
17/09/2026
CP Communications has been named Best of Florida for Audio-Visual Services, and its subsidiary Red House Streaming has been named Best Movie and Recording Studi...
17/09/2026
Scripps Sports and ION have announced a broadcast partnership for the Shriners Children's East-West Bowl, the nation's oldest college all-star football ...
17/09/2026
The 17th annual espnW: Women + Sports Summit, presented by Toyota, will take place October 14-16 at the Ojai Valley Inn in Ojai, California, with virtual attend...
17/09/2026
The Professional Audio Manufacturers Alliance (PAMA) and Shure Incorporated have...
17/09/2026
Lithuanian broadcaster LNK Group has completed a migration of playout and media management for its five national television networks - LNK, BTV, TV1, InfoTV, an...
17/09/2026
Net Insight has announced a pan-Asian live media network created with its regional partners, giving broadcasters, production companies, and media service provid...
17/09/2026
After a high school football injury, the Orlando native found a new path in video production, gaining experience on ESPN college football broadcasts, Tampa Bay ...
17/09/2026
Popular DAW gains free pitch-correction tool
Image-Line have teamed up with Antares to kit their popular DAW software out with a built-in pitch-correction p...
17/09/2026
Promises unprecedented realism, dynamic nuance and clarity
Gibson have teamed up with acoustic pickup experts Baggs, creating a next-generation' pickup...
17/09/2026
Two new libraries join percussion line-up
VSL (Vienna Symphonic Library) have just launched two new libraries that bring some intriguing new instruments int...
17/09/2026
Hardware & plug-in versions gain new features
Arturia have just released a free update that brings some significant new features to their MiniFreak synthesi...
17/09/2026
Respected journalist Catalina Fl rez to succeed Anton Enus as SBS World News pre...
17/09/2026
The Department of Sport, Arts and Culture (DSAC), in partnership with the Nation...
17/09/2026
Calrec brings leading audio solutions and long-term business value to IBC 2026 We're looking forward to meeting up with you all on 11-14th September in Amst...
17/09/2026
Streaming Content Ratings (SCR) Data Shifts to Daily Delivery Cadence, Mirroring...
17/09/2026
August brought further stabilization in total TV viewing time. Poles spent an average of 3 hours and 31 minutes a day watching video content on TV glass-just 2 ...
17/09/2026
Post-sports seasonal shift redistributes viewership across Poland, while streami...
17/09/2026
Television Content Analytics Pte Ltd (TVC), a Singapore-based deep-tech company specialising in sports technology, artificial intelligence and live broadcast in...
17/09/2026
Rise WIB, the global advocacy group championing gender diversity and career progression across the media technology industry, today announced the shortlist for ...
17/09/2026
RALEIGH, N.C. - Capitol Broadcasting Company (CBC) and Hurricanes Holdings annou...
17/09/2026
Lithuanian broadcaster deploys Playout X and Framelight X across five networks, combining on-premises software-defined playout with cloud-based disaster recover...
17/09/2026
Introducing Adobe Photoshop Elements & Premiere Elements 2027
Brie Clayton September 17, 2026
0 Comments
New tools make it easier than ever to enhance y...
17/09/2026
LA JOLLA, CA-While the brain orchestrates metabolic processes in the body, it al...
17/09/2026
New series of Aistear an Amhr in shares extraordinary stories behind Ireland'...
17/09/2026
The Late Late Show is back!
Liam Neeson, Siobh n McSweeney, Caitr ona Balfe a...
17/09/2026
The Neighbours Effect: Gen Z Brings Back Big Hair, Double Denim and 80s soap sty...
17/09/2026
The bold, new legal drama starring Dominic West and Sienna Miller launches on Sky and NOW on 2nd OctThursday 17 September 2026
Official trailer released for Sk...
17/09/2026
Every week we read headlines about new threats from AI. From taking over jobs to disrupting critical infrastructure, spreading misinformation or even destroying...
17/09/2026
Arqiva selected by SANZAAR to deliver global distribution of 2026 international ...
17/09/2026
CULVER CITY, CALIFORNIA This evening at the 78th Primetime Emmy Awards, Apple TV...
17/09/2026
A new creature-catching adventure is ready to stream from the cloud this week. Pawprint Studio's Aniimo arrives on GeForce NOW at launch, inviting gamers to...
17/09/2026
RT Statement:
RT is today confirming that its position regarding participation in the Eurovision Song Contest remains unchanged. RT will not participate in...
16/09/2026
Grass Valley has been recognized with a 2026 Technology & Engineering Emmy Award...
16/09/2026
Sports-media leaders explore how cloud, AI, and virtualization are transforming ...
16/09/2026
With a full crew on site in Springfield, deep access to players, and a growing d...
16/09/2026
John Wilson attends The History of Concrete premiere during the 2026 Sundance Film Festival at The Yarrow Theatre on January 22, 2026, in Park City, Utah. (Ph...
16/09/2026
Powerful soft synth receives free update
Wavea have just launched a new and improved version of their feature-packed soft synth, kitting it out with over 20...