![](//www.4rfv.com/images/no_image.jpg)
Indigenous languages are under threat. Some 3,000 - three-quarters of the total - could disappear before the end of the century, or one every two weeks, according to UNESCO.
As part of a movement to protect such languages, New Zealand's Te Hiku Media, a broadcaster focused on the M ori people's indigenous language known as te reo, is using trustworthy AI to help preserve and revitalize the tongue.
Using ethical, transparent methods of speech data collection and analysis to maintain data sovereignty for the M ori people, Te Hiku Media is developing automatic speech recognition (ASR) models for te reo, which is a Polynesian language.
Built using the open-source NVIDIA NeMo toolkit for ASR and NVIDIA A100 Tensor Core GPUs, the speech-to-text models transcribe te reo with 92% accuracy. It can also transcribe bilingual speech using English and te reo with 82% accuracy. They're pivotal tools, made by and for the M ori people, that are helping preserve and amplify their stories.
There's immense value in using NVIDIA's open-source technologies to build the tools we need to ultimately achieve our mission, which is the preservation, promotion and revitalization of te reo M ori, said Keoni Mahelona, chief technology officer at Te Hiku Media, who leads a team of data scientists and developers, as well as M ori language experts and data curators, working on the project.
We're also helping guide the industry on ethical ways of using data and technologies to ensure they're used for the empowerment of marginalized communities, added Mahelona, a Native Hawaiian now living in New Zealand.
Building a House of Speech' Te Hiku Media began more than three decades ago as a radio station aiming to ensure te reo had space on the airwaves. Over the years, the organization incorporated television broadcasting and, with the rise of the internet, it convened a meeting in 2013 with the community's elders to form a strategy for sharing content in the digital era.
The elders agreed that we should make the stories accessible online for our community members - rather than just keeping our archives on cassettes in boxes - but once we had that objective, the challenge was how to do this correctly, in alignment with our strong roots in valuing sovereignty, Mahelona said.
Instead of uploading its video and audio sources to popular, global platforms - which, in their terms and conditions of use, require signing over certain rights related to the content - Te Hiku Media decided to build its own content distribution platform.
Called Whare K rero - meaning house of speech - the platform now holds more than 30 years' worth of digitized, archival material featuring about 1,000 hours of te reo native speakers, some of whom were born in the late 19th century, as well as more recent content from second-language learners and bilingual M ori people.
Now, around 20 M ori radio stations use and upload their content to Whare K rero. Community members can access the content through an app.
It's an invaluable resource of acoustic data, Mahelona said.
Turning to Trustworthy AI Such a trove held incredible value for those working to revitalize the language, the Te Hiku Media team quickly realized, but manual transcription required pulling lots of time and effort from limited resources. So began the organization's trustworthy AI efforts, in 2016, to accelerate its work using ASR.
No one would have a clue that there are eight NVIDIA A100 GPUs in our derelict, rundown, musky-smelling building in the far north of New Zealand - training and building M ori language models, Mahelona said. But the work has been game-changing for us.
To collect speech data in a transparent, ethically compliant, community-oriented way, Te Hiku Media began by explaining its cause to elders, garnering their support and asking them to come to the station to read phrases aloud.
It was really important that we had the support of the elders and that we recorded their voices, because that's the sort of content we want to transcribe, Mahelona said. But eventually these efforts didn't scale - we needed second-language learners, kids, middle-aged people and a lot more speech data in general.
So, the organization ran a crowdsourcing campaign, K rero M ori, to collect highly labeled speech samples according to the Kaitiakitanga license, which ensures Te Hiku Media uses the data only for the benefit of the M ori people.
In just 10 days, more than 2,500 signed up to read 200,000+ phrases, providing over 300 hours of labeled speech data, which was used to build and train the te reo M ori ASR models.
In addition to other open-source trustworthy AI tools, Te Hiku Media now uses the NVIDIA NeMo toolkit's ASR module for speech AI throughout its entire pipeline. The NeMo toolkit comprises building blocks called neural modules and includes pretrained models for language model development.
It's been absolutely amazing - NVIDIA's open-source NeMo enabled our ASR models to be bilingual and added automatic punctuation to our transcriptions, Mahelona said.
Te Hiku Media's ASR models are the engines running behind Kaituhi, a te reo M ori transcription service now available online.
The efforts have spurred similar ASR projects now underway by Native Hawaiians and the Mohawk people in southeastern Canada.
It's indigenous-led work in trustworthy AI that's inspiring other indigenous groups to think: If they can do it, we can do it, too,' Mahelona said.
Learn more about NVIDIA-powered trustworthy AI, the NVIDIA NeMo toolkit and how it enabled a Telugu language speech AI breakthrough.
More from Nvidia
25/07/2024
Hey, you. You're finally awake.
It's the summer of Elder Scrolls - whe...
24/07/2024
AI is set to transform the workforce - and the Georgia Institute of Technology's new AI Makerspace is helping tens of thousands of students get ahead of the...
24/07/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
23/07/2024
Generative AI applications have little, or sometimes negative, value without acc...
23/07/2024
Businesses seeking to harness the power of AI need customized models tailored to their specific industry needs.
NVIDIA AI Foundry is a service that enables ent...
22/07/2024
AI and accelerated computing - twin engines NVIDIA continuously improves - are d...
22/07/2024
Team NVIDIA has triumphed at the Amazon KDD Cup 2024, securing first place Friday across all five competition tracks.
The team - consisting of NVIDIANs Ahmet E...
19/07/2024
Research published earlier this month in the science journal Nature used NVIDIA-powered supercomputers to validate a pathway toward the commercialization of qua...
19/07/2024
Research published earlier this month in the science journal Nature used NVIDIA-powered supercomputers to validate a pathway toward the commercialization of qua...
19/07/2024
AI has seen unprecedented growth - spurring the need for new training and educat...
19/07/2024
Research published earlier this month in the science journal Nature used NVIDIA-powered supercomputers to validate a pathway toward the commercialization of qua...
18/07/2024
Mistral AI and NVIDIA today released a new state-of-the-art language model, Mist...
18/07/2024
It's time for a sweet treat - the GeForce NOW Summer Sale offers high-perfor...
17/07/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
16/07/2024
Editor's note: This post is part of our In the NVIDIA Studio series, which c...
15/07/2024
NVIDIA founder and CEO Jensen Huang and Meta founder and CEO Mark Zuckerberg wil...
12/07/2024
NVIDIA is taking an array of advancements in rendering, simulation and generativ...
11/07/2024
Unlock new experiences every GFN Thursday. Whether post-apocalyptic survival adventures, narrative-driven games or vast, open worlds, GeForce NOW always has som...
11/07/2024
Enhancing Japan's AI sovereignty and strengthening its research and development capabilities, Japan's National Institute of Advanced Industrial Science ...
10/07/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible and showcases new hardware, softwar...
10/07/2024
Improved cancer diagnostics - and improved patient outcomes - could be among the...
09/07/2024
Sphere, a new kind of entertainment medium in Las Vegas, is joining the ranks of legendary circular performance spaces such as the Roman Colosseum and Shakespea...
08/07/2024
Artificial intelligence is transforming the transportation industry, helping dri...
04/07/2024
GeForce NOW is bringing 22 new games to members this month.
Dive into the four titles available to stream on the cloud gaming service this week to stay cool an...
03/07/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
28/06/2024
On a weekday afternoon, Ashwini Ashtankar sat on the bank of the Doodhpathri River, in a valley nestled in the Himalayas. Taking a deep breath, she noticed that...
27/06/2024
Editor's note: This post is part of Into the Omniverse, a series focused on ...
27/06/2024
Get ready to feel some chills, even amid the summer heat. Capcom's award-winning Resident Evil Village brings a touch of horror to the cloud this GFN Thursd...
26/06/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
26/06/2024
Roblox is a colorful online platform that aims to reimagine the way that people ...
25/06/2024
Generative AI has revolutionized software development with prompt-based code generation - protein design is next.
EvolutionaryScale today announced the release...
24/06/2024
Multi-die chips, known as three-dimensional integrated circuits, or 3D-ICs, represent a revolutionary step in semiconductor design. The chips are vertically sta...
20/06/2024
Sit back and settle in for some epic storytelling. Tell Me Why and As Dusk Falls - award-winning, narrative-driven games from Xbox Studios - add to the 1,900+ g...
19/06/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible and showcases new hardware, softwar...
18/06/2024
The electric grid and the utilities managing it have an important role to play in the next industrial revolution that's being driven by AI and accelerated c...
17/06/2024
NVIDIA contributed the largest ever indoor synthetic dataset to the Computer Vision and Pattern Recognition (CVPR) conference's annual AI City Challenge - h...
17/06/2024
Making moves to accelerate self-driving car development, NVIDIA was today named an Autonomous Grand Challenge winner at the Computer Vision and Pattern Recognit...
17/06/2024
NVIDIA researchers are at the forefront of the rapidly advancing field of visual...
15/06/2024
NVIDIA founder and CEO Jensen Huang on Friday encouraged Caltech graduates to pu...
14/06/2024
When she was five years old, Veronica Miller (n e Teklai) and her family left their homeland of Eritrea, in the Horn of Africa, to escape an ongoing war with Et...
14/06/2024
NVIDIA today announced Nemotron-4 340B, a family of open models that developers ...
13/06/2024
Set sail for adventure, pirates. Sea of Thieves makes waves in the cloud this week. It's an adventure-filled GFN Thursday with four new games joining the Ge...
12/06/2024
Accelerated computing is transforming data processing and analytics for enterpri...
12/06/2024
The full-stack NVIDIA accelerated computing platform has once again demonstrated...
12/06/2024
Let's talk about NeRFs - no, not the neon-colored foam dart blasters, but ne...
12/06/2024
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
07/06/2024
Across industries, AI is supercharging innovation with machine-powered computation. In finance, bankers are using AI to detect fraud more quickly and keep accou...
06/06/2024
Capcom's latest entry in the iconic Street Fighter series, Street Fighter 6, punches its way into the cloud this GFN Thursday. The game, along with Ubisoft&...
05/06/2024
India's AI market is expected to be massive. Yotta Data Services is setting its sights on supercharging it. In this episode of NVIDIA's AI Podcast, Suni...
05/06/2024
NVIDIA launched NVIDIA Studio at COMPUTEX in 2019. Five years and more than 500 ...