
Linguists estimate that at least half of the world's estimated 7,000 spoken languages will become extinct by the century's end, due to forces ranging from globalization to cultural assimilation.
Part of the challenge of documenting and revitalizing endangered languages is a lack of texts and speech recordings to work with. Seneca, a language of one of the six Iroquois Nations in North America, has only about 100 first-language speakers and several hundred more second-language learners.
Automatic speech recognition (ASR) technology is widely used to transcribe languages with millions or billions of speakers, like English and Mandarin. But it has only scratched the surface with languages like Seneca, which have vastly fewer speakers and significantly less data to work with.
Now a team of researchers at the Rochester Institute of Technology in New York, along with colleagues from the University at Buffalo, is tapping deep learning to bolster the ability of ASR. And while its focus is on Seneca, the project's vision encompasses the preservation of languages globally as well as an important part of our shared cultural history.
Knowing about different languages teaches us a lot about how our brain works, said Emily Prud'hommeaux, an assistant professor of computer science at Boston College and a research faculty member at RIT. When you document a language, you're preserving information not only about that language but also about how humans use language in general.
It's no coincidence that Prud'hommeaux and her team started with the Seneca language. Three members of the Seneca nation are part of the effort - a direct connection that is rare in research of this type, she said.
Leading the charge is Robbie Jimerson, a Ph.D. student in RIT's Golisano College of Computing and Information Science. He is a member of the Seneca Nation of Indians and is passionate about ensuring the survival of the Seneca language.
There's a big effort by the leaders of the tribe to preserve and promote our language, said Jimerson. I was looking for an opportunity to contribute.
Using GANs to Create More Language Samples Now in its third year, the project has had challenges when it comes to accumulating language data. Jimerson said the Seneca community can be guarded about what it shares with other people, so there wasn't an abundance of recordings of the language being spoken. He set out to change that.
He started by recording friends and elders who speak the language and asking them to record their friends. He found out whenever someone was speaking Seneca in public. He asked for family recordings of elders telling stories handed down from previous generations. And he grabbed any publicly available videos or recordings he could find online.
The team has fine-tuned an ASR model for Seneca, running it through generative adversarial networks to create more samples out of the limited number of recordings. The model turns wave files of the spoken language into streams of characters, while computing probability and making corrections.
The resulting data is fed into a deep learning model that in turn expands upon the ASR model's accuracy.
The team's networks run in two compute settings: on a nine-server machine learning lab running a variety of NVIDIA Tesla GPUs, and on a university cluster of large servers, each running 10 NVIDIA Tesla P4 GPUs. Each cluster runs a range of deep learning frameworks such as TensorFlow and Caffe.
The computer engineering cluster is for all students in the computer engineering department, and so they have to compete' for these resources, said Ray Ptucha, assistant professor of computer engineering at RIT, another collaborator on this project.
With access to these clusters at a premium, Jimerson tests code and checks the stability of models on a local machine running an NVIDIA TITAN X rather than inconvenience other students by running a model that might crash.
Achieving Better Accuracy So far, the team's efforts have brought the word error rate of its ASR model from 70 percent down to 56 percent. The goal, said Prud'hommeaux, is to get that rate down to 25 percent, which is where ASR systems were in processing English several years ago.
The more samples of spoken and written Seneca the team can accumulate, the more the error rate will decrease. (Today, English ASR models can achieve word error rates as low as 5 percent.)
The team's work is expected to help with language preservation efforts around the world.
Prud'hommeaux said the team has an agreement with an archiving institution that's a condition of a grant the project received from the National Science Foundation. The resulting language archiving database will be made available as a resource for other efforts seeking to document threatened languages.
Additionally, Prud'hommeaux said the team's work could prove helpful for any deep learning effort that has to make do with limited amounts of data.
Read more about the team's work in their research papers here and here.
Feature image: The Haudenosaunee (Iroquois Confederacy) flag, via Wikimedia Commons.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
10/09/2026
Tedial is introducing Media Business Hub (MBH) at IBC 2026, a new business enablement platform designed to help media organizations unlock the full value of the...
10/09/2026
Globecast, a leading provider of broadcast, media and entertainment managed services, will unveil the next phase of its transformation at IBC2026 with a preview...
10/09/2026
Fabric Data and TMT Insights today announced that the two companies are combining to create Fabric Insights, uniting authoritative data, purpose-built software,...
10/09/2026
Operative, the preferred advertising and broadcast management provider for the world's leading media brands, today announced a comprehensive slate of new pr...
10/09/2026
Test & measurement innovator, Leader Electronics of Europe, has announced the launch of its latest range of handheld signal generation, analysis and monitoring ...
10/09/2026
The partnership begins with native Zixi protocol support on the Kiloview D350 and will extend to deeper product integration and joint go-to-market initiatives
...
10/09/2026
Avid and FilmLight announce strategic partnership to reimagine the color pipelin...
10/09/2026
Maxon Brings Exclusive Preview of Fall Innovations to IBC2026
Brie Clayton September 10, 2026
0 Comments
IBC attendees will get the first look at new ...
10/09/2026
Puget Systems Launches Built Different by Puget Systems, A New Podcast That Go...
10/09/2026
Three Boston Conservatory Dance Alums Honored for Their Early-Career Promise India Hobbs and Demetrius Lee received Princess Grace Awards, while Lilly David t...
10/09/2026
MonoNeon Returns to Berklee to Kick Off Signature Series The Grammy-winning bassist and alumnus performs with students and faculty from the Planet MicroJam In...
10/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
10/09/2026
RT to Broadcast Amgen Irish Open free-to-air
New multi-year agreement will see RT continue to bring the Amgen Irish Open to Irish audiences
Ryder Cup hig...
10/09/2026
More than 30 women report allegations of sexual harassment, misogyny and bullying within National Ambulance Service
Tonight on RT Prime Time, 9:35pm, RT One ...
10/09/2026
Series 3 of Nero's Class begins Friday 11 September on RT Listen
The third...
10/09/2026
Gear up: The latest PC games and major updates are ready to play on GeForce NOW ...
10/09/2026
AI inference chipmaker d-Matrix today announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure pl...
09/09/2026
Sony Electronics has announced the SEL814G, the company's first fisheye zoom G Lens for full-frame Alpha E-mount cameras. The lens will be available in Octo...
09/09/2026
Telestream has announced an expansion of its Practical AI capabilities across th...
09/09/2026
Haivision and Chyron have announced a technology partnership integrating Haivisi...
09/09/2026
Omaha Productions has released ManningCast: Heist, a short film featuring Peyton and Eli Manning attempting to break into NFL Commissioner Roger Goodell's o...
09/09/2026
Imagine Communications will demonstrate the next evolution of XVR at IBC2026 (Stand 1.B73, RAI Amsterdam, September 11-14).
XVR is Imagine's channel-in-a-b...
09/09/2026
swXtch.io has announced abbe, an AI-based orchestration layer for live media workflows, making its world premiere at IBC2026 (Meeting Room 5.MR13, Hall 5, Septe...
09/09/2026
NFL Films has announced the second season of Inside the NFL on X as an X Original, premiering Wednesday, September 9. The series marks Inside the NFL's 50th...
09/09/2026
LiveU has announced an expanded partnership with The Associated Press, deploying the largest fleet of LU900Q intelligent production units in the industry as par...
09/09/2026
Diversified has announced plans to consolidate its four existing New Jersey locations into a new facility in Newark, expected to be fully operational by the end...
09/09/2026
Wisycom will exhibit new RF products at IBC2026 (Stand 8.D60, September 11-14), including the MATF Wideband Antenna Matrix, PFL Portable RF over Fiber Box, and ...
09/09/2026
Guntermann & Drunck (G&D) and VuWall, both Panoptec companies, will introduce vsNEO at IBC2026 (Stand 8.B63, September 11, RAI Amsterdam). Pre-orders open in ea...
09/09/2026
Mo-Sys has completed a business restructuring and relocated to a new facility in Canada Water, London. As part of the reorganization, Selin Kemal has been promo...
09/09/2026
Aicox has integrated TAG Video Systems' Core monitoring platform into Telef nica's TSAmediaHUB audiovisual contribution and distribution infrastructure....
09/09/2026
Blackmagic Design has released DaVinci Resolve 21.1, a free update available now through the Blackmagic Design website. The release adds new AI tools, workflow ...
09/09/2026
ESPN and Apple Music have announced a collaboration on music curation for Monday...
09/09/2026
Sacramento Production Services has completed a wireless audio overhaul at Sutter Health Park, home of the Athletics during their three-year residency in Sacrame...
09/09/2026
NDI and NVIDIA have announced a collaboration to develop AI-powered workflows enabling broadcasters to create and distribute multilingual content from a single ...
09/09/2026
Christopher Chip Adams, Troy Aikman, Jim Dove, Jim Eady, Bucky Gunts, Robert A. Iger, Charlie Jablonski, Robert Kraft, Tom Rinaldi, and Dick Stockton to be ho...
09/09/2026
ND2 returns as the primary unit, While FNIA gets SSQC and new flexible Rover...
09/09/2026
New parts aimed at hip-hop and pop productions
The latest update to Celemony's innovative virtual session musician expands the collection of built-in in...
09/09/2026
Save up to 40% until 21 September 2026
Sonarworks are currently running their Autumn Sale, with discounts of up to 40% currently available across their rang...
09/09/2026
Upcoming title officially launches in September 2026
Bjooks' celebrations for this year's 909 Day will be welcome news amongst drum machine lovers, ...
09/09/2026
NITV to bring First Nations netball to audiences nationwide
9 September, 2026
Media releases
Every match of the 2026 First Nations Tournament will be live ...
09/09/2026
The National Film and Video Foundation (NFVF) invites eligible Tier 1 and 2 South African production companies to submit proposals to serve as the Facilitating ...
09/09/2026
Italy launch follows UK and Germany as Nielsen expands product across EMEA
Milan, Italy, September 9, 2026 - Nielsen, a global leader in audience measurement, ...
09/09/2026
Planned solutions tackle accuracy, cost and latency barriers to deploying AI at ...
09/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...