Sony Pixel Power calrec Sony

How Deep Learning Is Aiding Preservation of Seneca and Other Endangered Languages

03/01/2019

Linguists estimate that at least half of the world's estimated 7,000 spoken languages will become extinct by the century's end, due to forces ranging from globalization to cultural assimilation.

Part of the challenge of documenting and revitalizing endangered languages is a lack of texts and speech recordings to work with. Seneca, a language of one of the six Iroquois Nations in North America, has only about 100 first-language speakers and several hundred more second-language learners.

Automatic speech recognition (ASR) technology is widely used to transcribe languages with millions or billions of speakers, like English and Mandarin. But it has only scratched the surface with languages like Seneca, which have vastly fewer speakers and significantly less data to work with.

Now a team of researchers at the Rochester Institute of Technology in New York, along with colleagues from the University at Buffalo, is tapping deep learning to bolster the ability of ASR. And while its focus is on Seneca, the project's vision encompasses the preservation of languages globally as well as an important part of our shared cultural history.

Knowing about different languages teaches us a lot about how our brain works, said Emily Prud'hommeaux, an assistant professor of computer science at Boston College and a research faculty member at RIT. When you document a language, you're preserving information not only about that language but also about how humans use language in general.

It's no coincidence that Prud'hommeaux and her team started with the Seneca language. Three members of the Seneca nation are part of the effort - a direct connection that is rare in research of this type, she said.

Leading the charge is Robbie Jimerson, a Ph.D. student in RIT's Golisano College of Computing and Information Science. He is a member of the Seneca Nation of Indians and is passionate about ensuring the survival of the Seneca language.

There's a big effort by the leaders of the tribe to preserve and promote our language, said Jimerson. I was looking for an opportunity to contribute.

Using GANs to Create More Language Samples Now in its third year, the project has had challenges when it comes to accumulating language data. Jimerson said the Seneca community can be guarded about what it shares with other people, so there wasn't an abundance of recordings of the language being spoken. He set out to change that.

He started by recording friends and elders who speak the language and asking them to record their friends. He found out whenever someone was speaking Seneca in public. He asked for family recordings of elders telling stories handed down from previous generations. And he grabbed any publicly available videos or recordings he could find online.

The team has fine-tuned an ASR model for Seneca, running it through generative adversarial networks to create more samples out of the limited number of recordings. The model turns wave files of the spoken language into streams of characters, while computing probability and making corrections.

The resulting data is fed into a deep learning model that in turn expands upon the ASR model's accuracy.

The team's networks run in two compute settings: on a nine-server machine learning lab running a variety of NVIDIA Tesla GPUs, and on a university cluster of large servers, each running 10 NVIDIA Tesla P4 GPUs. Each cluster runs a range of deep learning frameworks such as TensorFlow and Caffe.

The computer engineering cluster is for all students in the computer engineering department, and so they have to compete' for these resources, said Ray Ptucha, assistant professor of computer engineering at RIT, another collaborator on this project.

With access to these clusters at a premium, Jimerson tests code and checks the stability of models on a local machine running an NVIDIA TITAN X rather than inconvenience other students by running a model that might crash.

Achieving Better Accuracy So far, the team's efforts have brought the word error rate of its ASR model from 70 percent down to 56 percent. The goal, said Prud'hommeaux, is to get that rate down to 25 percent, which is where ASR systems were in processing English several years ago.

The more samples of spoken and written Seneca the team can accumulate, the more the error rate will decrease. (Today, English ASR models can achieve word error rates as low as 5 percent.)

The team's work is expected to help with language preservation efforts around the world.

Prud'hommeaux said the team has an agreement with an archiving institution that's a condition of a grant the project received from the National Science Foundation. The resulting language archiving database will be made available as a resource for other efforts seeking to document threatened languages.

Additionally, Prud'hommeaux said the team's work could prove helpful for any deep learning effort that has to make do with limited amounts of data.

Read more about the team's work in their research papers here and here.

Feature image: The Haudenosaunee (Iroquois Confederacy) flag, via Wikimedia Commons.
LINK: https://blogs.nvidia.com/blog/2019/01/02/deep-learning-preserves-senec...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

06/05/2026

Wisycom RF Solutions Support Gravity Medias Live Cycling and Marathon Broadcasts

Gravity Media Chief RF Communications Engineer Glenn Willems uses Wisycom RF over Fiber and wireless solutions across major cycling events and international mar...

06/05/2026

Sennheiser Spectera Module Now Available in Bitfocus Companion and Buttons

A Sennheiser Spectera module is now available in Bitfocus Companion and Buttons, enabling direct integration of Spectera with the two software platforms. The mo...

06/05/2026

Ted Turner, Cable Television Pioneer, Sports Broadcasting Hall of Famer, Dead at 87

Ted Turner, the visionary media entrepreneur whose appetite for disruption helpe...

06/05/2026

FIFA World Cup 2026: Peacock Launches Visin de Campo (aka Pitchside Live), Will Stream All 104 Matches in Spanish

Peacock is going all-in on the beautiful game - streaming all 104 FIFA World Cup...

06/05/2026

NoiseWorks Audio launch VoiceAssist Basic, Standard & Advanced

New pricing tiers for vocal/dialogue restoration tool NoiseWorks Audio's AI-powered vocal and dialogue processing plug-in is now available in three diff...

06/05/2026

RME TotalMix FX 2 now available

Popular mixing & routing software overhauled Following a recent public beta test, RME have launched the final release version of the powerful mixing and rou...

06/05/2026

Focusrite: Designing The ISA C8X Audio Interface

New SOS Video Feature Focusrites ISA C8X is a milestone product that brings together the companys analogue heritage and their expertise in digital audio. Yo...

06/05/2026

L3Harris to Boost Polish Navy Combat Power with Advanced Ship System

Polands Miecznik-class frigates are part of the largest contract in Polish shipbuilding history. (Image Credit: PGZ Stocznia Wojenna)...

06/05/2026

FCC's Anna Gomez Urges Rigorous Review of Paramount-WBD Merger

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Riedel Ups Marc Engroff to CFO, Shifts Frank Eischet to Group COO

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Amagi Launches In-Content Ads' to Attract More CTV Advertisers

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Riedel Expands Leadership Structure Appoints Marc Engroff...

Riedel Communications today announced the expansion of its leadership structure as part of a strategic initiative to strengthen both its operational management ...

06/05/2026

Production Sound Mixer Dirk Sciarrotta Delivers Camera Re...

For nearly three decades, Veteran Production Sound Mixer and Five-time Emmy Award Winner Dirk Sciarrotta has helped define the sonic identity of the long-runnin...

06/05/2026

ZEISS CinCraft LensCore: Cinema Lens Looks for Compositing

ZEISS CinCraft LensCore: Cinema Lens Looks for Compositing Brie Clayton May 6, 2026 0 Comments ZEISS announces the launch of CinCraft LensCore, a nove...

06/05/2026

Wisycom Solves Extreme RF Challenges Across Miles of Live Action for Gravity Media

Wisycom Solves Extreme RF Challenges Across Miles of Live Action for Gravity Med...

06/05/2026

NAB Launches Weekly Podcast on Local Broadcast Policy

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Mavis Launches Mavis Studio iPad For Media Production

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Narrative Entertainment partners with Encompass to provid...

Narrative Entertainment has partnered with Encompass to deliver high-quality subtitling of its Great! network content using the Altitude Intelligence AI assiste...

06/05/2026

SipRadius extends its seamless creation and connectivity...

SipRadius, widely recognized for making content processing and connectivity secure and seamless, is proud to launch a dramatic new approach to AI content creati...

06/05/2026

Big Blue Marble at ANGA COM - TV as a Service in the spot...

When the broadband and media industry gathers at ANGA COM in Cologne from May 19 to 21, Big Blue Marble will be at the forefront. The international broadcast an...

06/05/2026

Cinegy makes its MPTS debut with software-defined televis...

Cinegy GmbH, a leading developer of software-defined television technology, is proud to exhibit at MPTS for the first time. Visitors to the stand will discover ...

06/05/2026

Val Jeanty Receives 2026 Doris Duke Artist Award

Val Jeanty Receives 2026 Doris Duke Artist Award Jeanty, a composer, percussionist, and turntablist, is the fourth Berklee recipient of the prestigious award ...

06/05/2026

Zeiss Launches CinCraft LensCore

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Gomez Urges Rigorous FCC Review of Paramount-WBD Merger

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Wisycom Solves Extreme RF Challenges Across Miles of Live...

When live cycling races and international marathons stretch for miles across cities and countryside, there is no margin for RF failure in live broadcast. As Chi...

06/05/2026

ZEISS CinCraft LensCore - Cinema Lens Looks for Compositi...

Oberkochen/Germany, May 5, 2026 ZEISS announces the launch of CinCraft LensCore, a novel solution for creating physically based cinematic lens looks for visual...

06/05/2026

VEON's Kyivstar Authorized to Resell Starlink for Businesses & Enterprises in Ukraine

06 May 2026 VEON's Kyivstar Authorized to Resell Starlink for Businesses & ...

06/05/2026

UKTV Brings the Early Years of Neighbours to U in 420Episode Back Catalogue Deal

UKTV has secured the exclusive rights to the early back catalogue of iconic Australian drama series Neighbours, following a landmark content deal with Fremantle...

06/05/2026

Sky commissions feature documentary to mark 10th anniversary of Manchester Arena bomb

Wednesday 6 May 2026 Sky commissions feature documentary to mark 10th anniversa...

06/05/2026

Sky and Formula 1 agree long-term partnership across UK, Ireland and Italy

Wednesday 6 May 2026 Sky and Formula 1 agree long-term partnership across UK, Ireland and Italy Sky in the UK & Ireland to remain the home of Formula 1 until ...

06/05/2026

GUPPYFRIEND champions microplastic reduction in sport with The Real Bag For Life Sky TV campaign

Sky Zero Footprint Fund-backed TV campaign launches nationwide, supported by new...

06/05/2026

Riedel Expands Leadership Structure, Appoints Marc Engroff as CFO and Frank Eischet as Group COO

Wuppertal May 6, 2026 Riedel Expands Leadership Structure, Appoints Marc Engro...

06/05/2026

Harmonic Defines a New Era of Broadband Connectivity at ANGA COM 2026

SAN JOSE, Calif. - May 6, 2026 - Harmonic (NASDAQ: HLIT) today announced its latest broadband innovations for ANGA COM 2026, highlighting its vision for a new e...

06/05/2026

Statement from Rupert Murdoch, Chairman Emeritus, Fox Corporation on Ted Turner's Passing

Statement from Rupert Murdoch, Chairman Emeritus, Fox Corporation on Ted Turner&...

06/05/2026

Calling all animal lovers, The Shelter: Animal SOS returns for a sixth season

Friday 8 May on RT One and RT Player Meet the NSPCA team caring for and protecting animals in need in this six-part series Fly on the wall, six-part series...

06/05/2026

NVIDIA Spectrum-X - the Open, AI-Native Ethernet Fabric - Sets the Standard for Gigascale AI, Now With MRC

The race to build the world's most powerful AI factories demands networking ...

06/05/2026

May 05, 2026

How changes to proteins can alter drug interactions for new precision therapies Scripps Research team maps how chemical modifications to proteins affect drug bi...

05/05/2026

Lessons from fragile contexts on responding to disinformation

Experts from the world of academia, tech, business, politics and media convened for a Thomson Talks at the Cambridge Disinformation Summit in April. It's th...

05/05/2026

Samsung Galaxy S26 Ultra Phone Cameras Bring New Excitement to Street League Skateboarding

Three phones were hardwired for power and transmission to the truck; camera feat...

05/05/2026

Case Study: How Zaki Rose Rebuilt Its Production Infrastructure, and What It Means for Sports Content Creators

The creative studio behind campaigns for the NBA, Fanatics Sportsbook & Casino, ...

05/05/2026

Nielsen Co-Viewing Pilot Shows Average 4% Viewership Increase for February Live Events

Nielsen has announced results from a co-viewing pilot program covering February&...

05/05/2026

Nippon TV and FOR-A Win NAB Product of the Year and Future Best of Show Awards for viztrick AiDi

viztrick AiDi, an on-device AI solution developed by Nippon TV, delivered global...

05/05/2026

ARRI Introduces Omnibar LED Linear Fixture for Film, Live Entertainment, and Content Creation

ARRI has announced Omnibar, a battery-powered, IP65-rated multi-color LED linear...

05/05/2026

France Tlvisions Becomes First Broadcaster to Deploy Imagine Communications SNP-XS

Imagine Communications has announced that France T l visions is the first broadc...

05/05/2026

WNBA Announces Historic Canadian Media Rights Agreement with Bell Media

The Women's National Basketball Association (WNBA) and Bell Media today announced a multiyear agreement to broadcast and stream WNBA games in Canada beginni...