
More than 75 million people speak Telugu, predominantly in India's southern regions, making it one of the most widely spoken languages in the country.
Despite such prevalence, Telugu is considered a low-resource language when it comes to speech AI. This means there aren't enough hours' worth of speech datasets to easily and accurately create AI models for automatic speech recognition (ASR) in Telugu.
And that means billions of people are left out of using ASR to improve transcription, translation and additional speech AI applications in Telugu and other low-resource languages.
To build an ASR model for Telugu, the NVIDIA speech AI team turned to the NVIDIA NeMo framework for developing and training state-of-the-art conversational AI models. The model won first place in a competition conducted in October by IIIT-Hyderabad, one of India's most prestigious institutes for research and higher education.
NVIDIA placed first in accuracy for both tracks of the Telugu ASR Challenge, which was held in collaboration with the Technology Development for Indian Languages program and India's Ministry of Electronics and Information Technology as a part of its National Language Translation Mission.
For the closed track, participants had to use around 2,000 hours of a Telugu-only training dataset provided by the competition organizers. And for the open track, participants could use any datasets and pretrained AI models to build the Telugu ASR model.
NVIDIA NeMo-powered models topped the leaderboards with a word error rate of approximately 13% and 12% for the closed and open tracks, respectively, outperforming by a large margin all models built on popular ASR frameworks like ESPnet, Kaldi, SpeechBrain and others.
What sets NVIDIA NeMo apart is that we open source all of the models we have - so people can easily fine-tune the models and do transfer learning on them for their use cases, said Nithin Koluguri, a senior research scientist on the conversational AI team at NVIDIA. NeMo is also one of the only toolkits that supports scaling training to multi-GPU systems and multi-node clusters.
Building the Telugu ASR Model The first step in creating the award-winning model, Koluguri said, was to preprocess the data.
Koluguri and his colleague Megh Makwana, an applied deep learning solution architect manager at NVIDIA, removed invalid letters and punctuation marks from the speech dataset that was provided for the closed track of the competition.
Our biggest challenge was dealing with the noisy data, Koluguri said. This is when the audio and the transcript don't match - in this case you cannot guarantee the accuracy of the ground-truth transcript you're training on.
The team cleaned up the audio clips by cutting them to be less than 20 seconds, chopped out clips of less than 1 second and removed sentences with a greater-than-30 character rate, which measures characters spoken per second.
Makwana then used NeMo to train the ASR model for 160 epochs, or full cycles through the dataset, which had 120 million parameters.
For the competition's open track, the team used models pretrained with 36,000 hours of data on all 40 languages spoken in India. Fine-tuning this model for the Telugu language took around three days using an NVIDIA DGX system, according to Makwana.
Inference test results were then shared with the competition organizers. NVIDIA won with around 2% better word error rates than the second-place participant. This is a huge margin for speech AI, according to Koluguri.
The impact of ASR model development is very high, especially for low-resource languages, he added. If a company comes forward and sets a baseline model, as we did for this competition, people can build on top of it with the NeMo toolkit to make transcription, translation and other ASR applications more accessible for languages where speech AI is not yet prevalent.
NVIDIA Expands Speech AI for Low-Resource Languages ASR is gaining a lot of momentum in India majorly because it will allow digital platforms to onboard and engage with billions of citizens through voice-assistance services, Makwana said.
And the process for building the Telugu model, as outlined above, is a technique that can be replicated for any language.
Of around 7,000 world languages, 90% are considered to be low resource for speech AI - representing 3 billion speakers. This doesn't include dialects, pidgins and accents.
Open sourcing all of its models on the NeMo toolkit is one way NVIDIA is improving linguistic inclusion in the field of speech AI.
In addition, pretrained models for speech AI, as part of the NVIDIA Riva software development kit, are now available in 10 languages - with many additions planned for the future.
And NVIDIA last month hosted its inaugural Speech AI Summit, featuring speakers from Google, Meta, Mozilla Common Voice and more. Learn more about Unlocking Speech AI Technology for Global Language Users by watching the presentation on demand.
Get started building and training state-of-the-art conversational AI models with NVIDIA NeMo.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
27/06/2026
There's no doubt that you've seen the world through Amy Vincent's ey...
27/06/2026
Brings together saturation & lo-fi effects
Following on from the release of their Voxcraft vocal-processing plug-in, UJAM have announced the launch of Retro...
27/06/2026
A record 4.84 million Australians choose SBS as the Socceroos advance at FIFA Wo...
27/06/2026
Why CRAS Upgraded to Symphony I/O MK II When an audio school runs studios all day, every day, gear doesn't just need to sound good , it needs to survive rea...
27/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/06/2026
Krotos Video to Sound Plugin Now Available for Adobe Premiere Pro
Brie Clayton June 26, 2026
0 Comments
Editors can analyze footage, generate synchron...
27/06/2026
Mirai Media Elevates Digital and Broadcast Productions with Blackmagic Design
Brie Clayton June 26, 2026
0 Comments
Studio uses Ultimatte 12 HD and Po...
27/06/2026
DURHAM, N.C. - JUNE 26, 2026 - Lutra Cafe & Bakery has opened its first brick-and-mortar location at American Tobacco Campus after owner Chris McLaurin operated...
26/06/2026
In-venue and creative video staffers at the professional and collegiate level ha...
26/06/2026
Strike Fighter League (SFL), a professional air combat digital sport combining f...
26/06/2026
Wisycom has announced three new additions to its professional wireless ecosystem...
26/06/2026
Eurovision Services inaugurated an expanded Master Control Room (MCR) in Madrid on June 1, 2026, building on a broadcast hub the company has operated in the cit...
26/06/2026
Midco Sports and the University of North Dakota (UND) have announced a two-year ...
26/06/2026
Guntermann and Drunck (G&D) and VuWall, both part of the Panoptec Technologies Group, have appointed Vutec (Pty) Ltd as exclusive distributor for their KVM and ...
26/06/2026
Visit Seattle, the official destination marketing organization for Seattle and King County, has launched what it describes as the world's first drone scoreb...
26/06/2026
CP Communications provided RF video, audio, and crew communications support for ...
26/06/2026
Produced by longtime partner Echo Entertainment, the action-sports property is now a team-based year-round league
The inaugural season of the MoonPay X Games L...
26/06/2026
The deal establishes MultiDyne Robotics and Motion Control, maintaining the well-known MRMC brand.MultiDyne Video & Fiber Optic Systems has acquired the assets ...
26/06/2026
PX1 will debut at Sonoma as TNT leans into super-slo-mo, drones, SMT data integr...
26/06/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
26/06/2026
Virtual session musician plug-in gains new percussion options
Celemony's latest update for their virtual session musician platform complements the exist...
26/06/2026
Half-size model joins Console 1 line-up
Shortly after the release of their new Flow Studio controller, Softube have announced the launch of another new surf...
26/06/2026
ELT Group and Rohde & Schwarz sign a cooperation agreement to explore commercial...
26/06/2026
For Teddy Swims sold-out I've Tried Everything But Therapy tour, event technology specialists, PRG, provided video, automation and lighting across 19 date...
26/06/2026
Modern exhibition and event venues face the challenge of seamlessly integrating traditional conference technology, professional broadcast workflows and IP-based...
26/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/06/2026
Neko Oji: The Guy That Got Reincarnated as a Cat Edited with DaVinci Resolve Stu...
26/06/2026
Adobe to Acquire Topaz Labs
Brie Clayton June 25, 2026
0 Comments
Adobe has seen strong demand for its AI products for creatives, including Adobe Fire...
26/06/2026
Berklee Students Earn Dedicated Section at Raindance Film Festival in London Five documentary short films produced in the Africana Studies Department screen a...
26/06/2026
How IMS Productions and FOX Sports scaled coverage of the 109th Indianapolis 500.
The last lap of this year's Indianapolis 500 delivered the kind of ending...
26/06/2026
Flicker Productions to produce five-part docu-reality series following women who have fallen for men in prison and have become TikTok sensations, with brands an...
26/06/2026
Catch up on the latest developments across Baselight and Daylight v7, Nara and F...
26/06/2026
26. June 2026 News
DFT is pleased to announce that a second Polar HQ film s...
26/06/2026
New documentary Freedom Founder: Thomas McKean and the American Revolution airs ...
25/06/2026
Launching a Career in Broadcast Engineering: Academic Paths and Essential Certif...
25/06/2026
This superstar shooter/storyteller from Central Indiana hopes to make his mark in the blossoming sports-documentary and -features space
In the live-sports-vid...
25/06/2026
Presidio and the National Hockey League have announced a multiyear renewal of their North American partnership. Presidio will remain an Official Technology Inno...
25/06/2026
Strike Fighter League (SFL) is the world's first professional air combat digital sport that combines elite human performance and physical immersion with cut...
25/06/2026
Rise, the award-winning advocacy group for gender diversity in the broadcast and media technology sector, is pleased to announce the global mentoring cohort for...
25/06/2026
The 2026 American Association of Professional Baseball (AAPB) All-Star Game will...
25/06/2026
Mediaproxy has named Heartland Video Systems (HVS) as its exclusive partner for US television broadcasting. The Wisconsin-based systems integrator will represen...
25/06/2026
Backblaze has formed an agreement with CoreWeave to create The Essential Cloud for AI.
Under the multi-exabyte, $335 million agreement, Backblaze will provide...