Sony Pixel Power calrec Sony

Mori Speech AI Model Helps Preserve and Promote New Zealand Indigenous Language

16/01/2024

Indigenous languages are under threat. Some 3,000 - three-quarters of the total - could disappear before the end of the century, or one every two weeks, according to UNESCO.

As part of a movement to protect such languages, New Zealand's Te Hiku Media, a broadcaster focused on the M ori people's indigenous language known as te reo, is using trustworthy AI to help preserve and revitalize the tongue.

Using ethical, transparent methods of speech data collection and analysis to maintain data sovereignty for the M ori people, Te Hiku Media is developing automatic speech recognition (ASR) models for te reo, which is a Polynesian language.

Built using the open-source NVIDIA NeMo toolkit for ASR and NVIDIA A100 Tensor Core GPUs, the speech-to-text models transcribe te reo with 92% accuracy. It can also transcribe bilingual speech using English and te reo with 82% accuracy. They're pivotal tools, made by and for the M ori people, that are helping preserve and amplify their stories.

There's immense value in using NVIDIA's open-source technologies to build the tools we need to ultimately achieve our mission, which is the preservation, promotion and revitalization of te reo M ori, said Keoni Mahelona, chief technology officer at Te Hiku Media, who leads a team of data scientists and developers, as well as M ori language experts and data curators, working on the project.

We're also helping guide the industry on ethical ways of using data and technologies to ensure they're used for the empowerment of marginalized communities, added Mahelona, a Native Hawaiian now living in New Zealand.

Building a House of Speech' Te Hiku Media began more than three decades ago as a radio station aiming to ensure te reo had space on the airwaves. Over the years, the organization incorporated television broadcasting and, with the rise of the internet, it convened a meeting in 2013 with the community's elders to form a strategy for sharing content in the digital era.

The elders agreed that we should make the stories accessible online for our community members - rather than just keeping our archives on cassettes in boxes - but once we had that objective, the challenge was how to do this correctly, in alignment with our strong roots in valuing sovereignty, Mahelona said.

Instead of uploading its video and audio sources to popular, global platforms - which, in their terms and conditions of use, require signing over certain rights related to the content - Te Hiku Media decided to build its own content distribution platform.

Called Whare K rero - meaning house of speech - the platform now holds more than 30 years' worth of digitized, archival material featuring about 1,000 hours of te reo native speakers, some of whom were born in the late 19th century, as well as more recent content from second-language learners and bilingual M ori people.

Now, around 20 M ori radio stations use and upload their content to Whare K rero. Community members can access the content through an app.

It's an invaluable resource of acoustic data, Mahelona said.

Turning to Trustworthy AI Such a trove held incredible value for those working to revitalize the language, the Te Hiku Media team quickly realized, but manual transcription required pulling lots of time and effort from limited resources. So began the organization's trustworthy AI efforts, in 2016, to accelerate its work using ASR.

No one would have a clue that there are eight NVIDIA A100 GPUs in our derelict, rundown, musky-smelling building in the far north of New Zealand - training and building M ori language models, Mahelona said. But the work has been game-changing for us.

To collect speech data in a transparent, ethically compliant, community-oriented way, Te Hiku Media began by explaining its cause to elders, garnering their support and asking them to come to the station to read phrases aloud.

It was really important that we had the support of the elders and that we recorded their voices, because that's the sort of content we want to transcribe, Mahelona said. But eventually these efforts didn't scale - we needed second-language learners, kids, middle-aged people and a lot more speech data in general.

So, the organization ran a crowdsourcing campaign, K rero M ori, to collect highly labeled speech samples according to the Kaitiakitanga license, which ensures Te Hiku Media uses the data only for the benefit of the M ori people.

In just 10 days, more than 2,500 signed up to read 200,000+ phrases, providing over 300 hours of labeled speech data, which was used to build and train the te reo M ori ASR models.

In addition to other open-source trustworthy AI tools, Te Hiku Media now uses the NVIDIA NeMo toolkit's ASR module for speech AI throughout its entire pipeline. The NeMo toolkit comprises building blocks called neural modules and includes pretrained models for language model development.

It's been absolutely amazing - NVIDIA's open-source NeMo enabled our ASR models to be bilingual and added automatic punctuation to our transcriptions, Mahelona said.

Te Hiku Media's ASR models are the engines running behind Kaituhi, a te reo M ori transcription service now available online.

The efforts have spurred similar ASR projects now underway by Native Hawaiians and the Mohawk people in southeastern Canada.

It's indigenous-led work in trustworthy AI that's inspiring other indigenous groups to think: If they can do it, we can do it, too,' Mahelona said.

Learn more about NVIDIA-powered trustworthy AI, the NVIDIA NeMo toolkit and how it enabled a Telugu language speech AI breakthrough.
LINK: https://blogs.nvidia.com/blog/te-hiku-media-maori-speech-ai/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

19/06/2026

IBC Show to Increase Focus on Networking, Startups

Share Copy link Facebook X Linkedin Bluesky Email...

19/06/2026

Irdeto Taps Axel Gallant as CEO

Share Copy link Facebook X Linkedin Bluesky Email...

19/06/2026

SMPTE Makes Its Standards Freely Accessible - Opening St...

SMPTE , the home of media professionals, technologists and engineers, has announced that its entire Standards catalog is now freely available to the global medi...

19/06/2026

nsign and BrightSign partner to expand deployment options...

nsign, the digital signage SaaS platform built around its core principle of Simplify Complexity, has announced a partnership with BrightSign , expanding the dep...

19/06/2026

Visual Productions Unveils RdmRelay2 Four-channel Relay...

Visual Productions announces the availability of its new RdmRelay2 at InfoComm 2026 (ACT Entertainment, Booth N6813). A networked, four-channel DMX relay, it is...

19/06/2026

Adobe Unveils Major Expansion of Creative Agent Across Firefly and Creative Cloud Apps Including Photoshop and Premiere

Adobe Unveils Major Expansion of Creative Agent Across Firefly and Creative Clou...

19/06/2026

Immersive Studio Metaverse Stage Innovates Storytelling with URSA Cine Immersive

Immersive Studio Metaverse Stage Innovates Storytelling with URSA Cine Immersive Brie Clayton June 18, 2026 0 Comments Two new narrative short films c...

19/06/2026

June 18, 2026

Lab studies explain how new cancer drug works as it enters patient testing Immunologists at Scripps Research show how a new, experimental drug revives immune ce...

18/06/2026

Ratings Roundup: Knicks-Spurs NBA Finals Is Most Watched Since Jordan Bulls Era; FIFA World Cup Opens Big

Ratings Roundup is a rundown of recent rating news and is derived from press rel...

18/06/2026

InfoComm 2026: PTZOptics Debuts Healthcare Visual Reasoning Integration with LayerJot

PTZOptics has unveiled new Visual Reasoning demonstrations at InfoComm 2026 (Boo...

18/06/2026

IBC 2026: Announces New Future Tech Ignite Initiative, Expanded Exhibitor Participation

IBC2026 will take place at the RAI Amsterdam from September 11-14, bringing toge...

18/06/2026

InfoComm 2026 Holds Inaugural Media Day with Product Announcements from 19 Exhibitors

InfoComm 2026 held its first-ever Media Day on June 17, providing journalists an...

18/06/2026

Info Comm 2026: FOR-A America Announces TAA-Compliant LED Display Package

FOR-A America has announced a Trade Agreements Act (TAA)-compliant LED display solution combining Alfalite's Litepix LED displays and Brompton Technology...

18/06/2026

SVG GameDay, Ep. 20: Cleveland Browns Kyle Millen - Run of Shows & Rock n Roll

In-venue and creative video staffers at the professional and collegiate level have one major thing in common: the intensity and attention to detail ramps up dur...

18/06/2026

International Federation of American Football, TMRW Sports Partner on Global Growth of Flag Football

The International Federation of American Football (IFAF) and TMRW Sports have an...

18/06/2026

AJA Debuts Io Xpand Thunderbolt 5 Expansion Chassis for KONA and Corvid PCIe I/O Cards

AJA Video Systems has unveiled Io Xpand, a Thunderbolt 5-enabled PCIe expansion ...

18/06/2026

ESPN Marks 30th Anniversary of WNBA's Inaugural Game With Liberty-Sparks Broadcast

ESPN has announced its coverage plans for the 30th anniversary of the WNBA's...

18/06/2026

FOX Sports' Big Noon Kickoff Heads to London for Union Jack Classic From Wembley Stadium

FOX Sports' Big Noon Kickoff will broadcast live from Wembley Stadium in Lon...

18/06/2026

InfoComm 2026: Show Opens With Microsoft Keynote, Media Day, and New Industry Initiatives

InfoComm 2026 opened on Wednesday at the Las Vegas Convention Center, bringing t...

18/06/2026

SVG New Sponsor Spotlight: Akta's Matt Smith on Building AI Into the Foundation of the Video Workflow

As media companies look to deliver more live, VOD, and snackable sports content ...

18/06/2026

2026 Sundance Film Festival: Local Lens

Top L-R: Take Me Home, The Lake Bottom L-R: TheyDream, Union County Free Summer Screening Series Announced Screenings for the Local Utah Community at...

18/06/2026

iamReverb gets an update

Improvements & new IR content iamReverb Audio have just launched a free update that kits their convolution reverb plug-in out with some new features and int...

18/06/2026

Two notes unveil Genome 2.0

Modelling suite gains improved captures, iOS support & more Two notes Audio Engineering have just announced the launch of Genome 2.0, a significant update t...

18/06/2026

VSL update Vienna Ensemble Pro 8

New AI assistance feature, video overhaul & more VSL have just announced the launch of Vienna Ensemble Pro 8.1 and 8.1V, a pair of major updates to their ev...

18/06/2026

The Gauge: Poland | May 2026

The average daily TV screen time in May was 3 hours and 36 minutes, marking a clear decrease of 15 minutes compared to April. This trend proved to be much stron...

18/06/2026

Gracenote and PubMatic Bring Curated Live Sports and Content-Level Deals to Programmatic CTV

Integration embeds Gracenote content intelligence, including contextual segments...

18/06/2026

AJA Debuts Io Xpand Thunderbolt 5 Expansion Chassis for...

AJA Video Systems unveiled Io Xpand, a high-performance Thunderbolt 5-enabled PCIe expansion chassis for AJA KONA and Corvid video and audio I/O cards. As dema...

18/06/2026

Providius Announces Providius Direct for Faster AV-over-I...

New NVRT-driven workflow enables on-demand traffic mirroring and guided troubleshooting for AV teams, integrators, and support organizations ahead of InfoComm 2...

18/06/2026

Mavis Studio Brings New Muscle to Mobile Production on iP...

At InfoComm 2026, Mavis announced a major update to Mavis Studio, its live production app for the iPad, with new features designed to make professional AV produ...

18/06/2026

Harmonic Completes Divestiture of Video Business to Media...

Transaction Positions Harmonic as a Pure-Play Broadband Company Harmonic Inc. (NASDAQ: HLIT), the worldwide leader in virtualized broadband solutions, today a...

18/06/2026

ACT Entertainment and NETGEAR Partner to Power the Futur...

ACT Entertainment and NETGEAR have announced a strategic partnership that establishes verified interoperability between NETGEAR Switches and key technologies fr...

18/06/2026

IBC2026 creates new pathways from innovation to real-worl...

IBC2026 is set to bring the global media, entertainment and technology community together at the RAI Amsterdam from 11 14 September, enabling industry players f...

18/06/2026

PLIANT TECHNOLOGIES DEBUTS CREWCOM FLEX THE WORLDS FIRST...

Pliant Technologies is proud to introduce CrewCom Flex , the world's first frameless matrix intercom system and the only complete matrix IP-based intercom s...

18/06/2026

XR Sports Alliance Strengthens its Capabilities with New...

The XR Sports Alliance (XRSA) has announced that a new cohort of members has joined the strategic initiative: ActionStreamer, Antigravity, Creative Artists Agen...

18/06/2026

Hollywood Filmmakers Launch AI Production Platform Cascade

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

ATVA Blasts Deltavision Media for Demanding 'Egregious' Retrans Fees

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

MediaKind Completes Merger with Harmonic's Video Business

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

AJA Unveils Io Xpand Thunderbolt 5 Expansion Chassis At InfoComm 2026

Share Copy link Facebook X Linkedin Bluesky Email...

18/06/2026

Linda May Han Oh Receives Guggenheim Fellowship

Linda May Han Oh Receives Guggenheim Fellowship The Berklee professor, bassist, and composer will use the fellowship to debut Dreams of Knowing, an interdisci...

18/06/2026

How FERC's Large-Load Interconnection Actions Help Address Grid Stress, Improve Affordability

In a consequential grid infrastructure decision, the Federal Energy Regulatory C...

18/06/2026

Techtel delivers Haivision live streaming system for TAFE NSW St Leonards, enabling real-world broadcast training

Techtel delivers Haivision live streaming system for TAFE NSW St Leonards, enabl...

18/06/2026

The Great Roaming Rinse

New study reveals 10 hidden data drainers costing Brits hundreds abroad - and the holiday hotspots where you could get rinsed the mostThursday 18 June 2026 The...

18/06/2026

How to watch the 2026/27 Scottish Premiership season on Sky Sports

Thursday 18 June 2026 How to watch the 2026/27 Scottish Premiership season on Sky Sports Which matches are Sky Sports showing on the 2026/27 Scottish Premiers...

18/06/2026

Sky Sale: Latest deals now on, with discounts on iPhone Air & 2.5Gbps speeds

Thursday 18 June 2026 Sky Sale: Latest deals now on, with discounts on iPhone Air & 2.5Gbps speeds The latest deals have dropped from Sky Mobile, the award-wi...

18/06/2026

FOX Advertising and Toonstar Team to Create New Opportunities for Brands in Digital-First Animation

FOX Advertising and Toonstar Team to Create New Opportunities for Brands in Digi...

18/06/2026

Arqiva secures WTA Tier-4 accreditation

Arqiva's Crawley Court and Chalfont Grove teleports re-certified at highest World Teleport Association standard 18 June 2026, Winchester, UK - Arqiva, the ...