Sony Pixel Power calrec Sony

Transcribing the future

11/10/2018

-- --

Facebook

Twitter

Google

Pinterest

SCREEN AFRICA EXCLUSIVE: A while back I was handed a bunch of very long audio files and was tasked with cutting about 60 hours of interviews, sound bites and voice over down to a 30-minute radio piece. Easy I thought, as long as I could get the material transcribed quickly and at reasonable cost and there my journey of discovery began.

Voice to text transcriptions have long been used in the media, medical and legal industries, traditionally done by human transcription teams. It's big business, but turn arounds can be slow and files often need a second error check to make sure the content is accurate. The cost of transcription actually wasn't a factor in my case - it was speed that I needed, I needed a machine to plough through my audio files and spit out a transcript so that I could search for key words and edit my story together.

-- --

Trolling the internet, I instantly found a few options and uploaded the same test file to all of them (as part of my free trials) but had really disappointing results ranging from 25 to about 59 per cent accuracy. Simple words and phrases were being interpreted as something completely different, more complex words like Fakarava Atoll came back as expletives! At first I thought that as the majority of interviews were heavy New Zealand accents that might be the problem but a snippet of the best British accented guest gave me similar results.

Through my work in the video world, I am aware that there is a lot of research and development in the transcription arena utilising Artificial Intelligence (AI). Machine learning works best when it is processing large analysable data sets like text. But most of the data being produced in the world right now is not text, it's the spoken word embedded in video and audio recording and thus the goal for AI developers to produce a reliable voice transcription process has intensified.

Tech companies like Apple, Google, Microsoft and Amazon are all actively involved in this space and have been researching voice recognition since the 90s and that research has only accelerated with the emergence of virtual assistants like Alexa, Cortana, and Google Voice and Siri. However, most people who use Siri or Alexa would agree that, while those tools do a reasonable job of understanding you, most of us wouldn't trust them with our lives. I asked Alexa where the Fakarava Atoll was and her response was, I would rather not answer that question. (Out of interest it is in Tahiti and is not a swear word!) A voice assistant like Alexa only needs to work out which, of a predetermined list of vocal commands is being asked, whereas a transcription programme needs to listen for and capture all spoken words and this wider variety of possible inputs and outputs makes it a more difficult task for AI.

Whilst stumbling around for my transcription answers I came across an article published by a team of Data Scientists and enthusiastic entrepreneurs, Ashutosh Trivedi and Anup Gosavi, who recently founded a company called Spext. Trivedi, based in Bangalore India, has deep interest and post graduate expertise in AI and has published his research in many IEEE journals. Gosavi is based in San Francisco and specialises in Design Thinking and Information Visualisation.

Spext describes their company name as a fusion of the words speech and text, and from the outset they looked like they could offer me exactly what I wanted and more. The service can best be described as a combined voice transcriber and media editor. You upload your audio files and the system automatically converts voice to text and displays the result in an edit window where it aligns the audio content with the text accurately and that means you can now do some amazing things with the resultant files. Not only do you get a full transcription of your work but you can edit the transcript, like you would on a word processor and then export the result as a new audio file. Obviously you can't create new sentences but the ability to edit and output the existing data as an audio file is a huge plus. It looks like a normal text editor and has familiar actions like copy-paste, cut-paste and I found editing by transcriptions on the fly extremely easy. When you are done you export your work as a word document, pdf and/or a new mp3 or wave file or even export your project to professional editing tools such as Adobe Audition, and Final Cut Pro for fine tuning.

The most important result is that the files I used to test other systems uploaded into the Spext system quickly and came back blazingly fast with a resultant accuracy of 96 per cent in my case. The system had even correctly punctuated the transcription, coped well with proper nouns like the names of fish species and fishing techniques but it too also battled with transcribing Fakarava (expletive) but at least recognised the word Atoll. It took no time at all to quickly manually edit any corrections. What could have taken weeks in production will easily get done in a matter of days now, artificial intelligence seems to have finally reached the point where transcription of audio by a machine works efficiently enough to make it viable and as researchers and companies improve and refine their algorithms, it seems evident that transcriptions will become even more accurate and the potential productivity savings of automated transcription will be hard to ignore. Someday soon, we might even be headed towards a world where audio files and text are thought of not as two distinct media types, but as two formats for the same content - as interchangeable - and as convertible as an .mp3 and a .wav file or a text file and a Word document.

The guys at Spext have used a unique combination of intuitive user experience design, to make it easy for the user, and advanced machine learning tha
LINK: http://www.screenafrica.com/2018/10/11/technology/ai-artificial-intell...
See more stories from screenafrica

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

06/09/2026

Dolby and MagentaTV Bring Fans Closer to the FIFA World Cup 2026 in Germany with Dolby Vision and Dolby Atmos

June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/07/2026

Adder Technology Names Neil Hillier as CEO

Share Copy link Facebook X Linkedin Bluesky Email...

01/07/2026

IBCAP Opens New Anti-Piracy Lab in Denver

Share Copy link Facebook X Linkedin Bluesky Email...

01/07/2026

FCC Plans to Auction 160 MHZ of Mid-Band Spectrum

Share Copy link Facebook X Linkedin Bluesky Email...

01/07/2026

CBS Miami Launches 'Hope 4 Venezuela' Relief Effort

Share Copy link Facebook X Linkedin Bluesky Email...

01/07/2026

Groundbreaking First Nations Screen Business Accelerator launched through national partnership

Groundbreaking First Nations Screen Business Accelerator launched through nation...

01/07/2026

Chyron Launches the All-New Chyron Academy: A Reimagined, Hands-On Learning Experience for Live Broadcast Production

Chyron Launches the All-New Chyron Academy: A Reimagined, Hands-On Learning Expe...

01/07/2026

Amplium Captures Kawasaki Brave Thunders Game with Blackmagic URSA Cine Immersive

Amplium Captures Kawasaki Brave Thunders Game with Blackmagic URSA Cine Immersiv...

01/07/2026

Boris FX Optics Expands Plugin Support to Apple Photos, Capture One, and Affinity Photo

Boris FX Optics Expands Plugin Support to Apple Photos, Capture One, and Affinit...

30/06/2026

Entries open for Thomson's Young Journalist Award 2026

Could your journalism reach an international stage? Entries are now open for the Thomson Foundation's Young Journalist Award 2026, one of the most prestigi...

30/06/2026

CazTVs 12 ENG Teams Across North America Keep Brazilian Fans on Top of World Cup

As Brazil's only way for fans to see all 104 matches, YouTube channel proves the power of digital...

30/06/2026

UJAM release Retrocraft multi-effects

Brings together saturation & lo-fi effects Following on from the release of their Voxcraft vocal-processing plug-in, UJAM have announced the launch of Retro...

30/06/2026

Zensphere v2 from Rapid Flow

New IR reverb engine, Juno-inspired chorus & more The latest version of Rapid Flow's hardware-emulation synth plug-in expands on its predecessor with a ...

30/06/2026

Shy Audio release Shy 90s Smack

Excels at heavy-handed VCA compression For their latest release, Shy Audio have recreated the crunchy' sound of a rackmount compressor that found its w...

30/06/2026

Apple raise Mac & iPad prices

Component scarcity drives cost increases Shortly after Apple's CEO Tim Cook acknowledged that cost increases would soon be inevitable , the company hav...

30/06/2026

Statement regarding GetUp Save Our SBS' campaign

Statement regarding GetUp Save Our SBS' campaign 30 June, 2026 Media releases The GetUp Save Our SBS' campaign is an independent initiative. SBS ...

30/06/2026

The First Hitachi Cash Recycling Devices in the EU Were Deployed at Bank Pekao S.A.

Hitachi and Bank Pekao S.A. have completed the installation of the first Hitachi...

30/06/2026

Clear-Com Upgrades Communication Systems for Jeopardy! and Wheel of Fortune

eds3_5_jq(document).ready(function($) { $(#eds_sliderM519).chameleonSlider_2_1({ content_source:......

30/06/2026

Telemundo, Peacock See More Record Setting World Cup Audiences

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

FOR-A America Adds Two Execs to U.S. Sales Team

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

A3SA Disputes Weigel Assertions that NextGen TV Threatens EAS

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

MainStreaming Selected by ITV to Support Delivery of ITVX...

MainStreaming, the award-winning and innovative Edge Video Delivery Network, today announced that it has been selected by ITV to support the delivery of ITVX, I...

30/06/2026

Chyron Launches New Chyron Academy

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

Clear-Com Upgrades Communication Systems for Jeopardy and...

When Wheel of Fortune and Jeopardy! needed to upgrade their wireless communications system, they turned to Clear-Com FreeSpeak wireless for their iconic televi...

30/06/2026

Supreme Court Gives Trump Tight Control over Independent Regulators

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

Rocket Lab to Acquire Iridium in $8 Billion Deal

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

Kyocera AVX Releases New Web-Based Antenna Integration Tool

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

YouTube Shorts Get a Makeover

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

Rise Announces 2026 Worldwide Mentoring Cohorts

Share Copy link Facebook X Linkedin Bluesky Email...

30/06/2026

Other World Computing Launches New Atlas Core Line with 256GB CFExpress 4.0 Type B Memory Card

Other World Computing Launches New Atlas Core Line with 256GB CFExpress 4.0 Type...

30/06/2026

DaVinci Resolve Studio Used for Taketoshi Sado's Perfume Cold Sleep -25 years Document-

DaVinci Resolve Studio Used for Taketoshi Sado's Perfume Cold Sleep -25 year...

30/06/2026

NVIDIA BioNeMo Agent Toolkit Brings Accelerated AI to Life Sciences Researchers in Claude Science

Life sciences has entered an era of computational scale, and for more than a dec...

30/06/2026

FOR-A America Expands U.S. Sales Team to Accelerate Growth of Software-Defined Solutions

Fernando Cruz and Jaz Wray Join as Regional Sales Managers; Bringing Extensive S...

30/06/2026

How NVIDIA's Inference Software Stack Powers the Lowest Token Cost

As organizations move from AI pilots to production AI factories, infrastructure decisions have shifted from peak chip specifications to cost per token: how many...

30/06/2026

Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning

Editor's note: This post is part of Into the Omniverse, a series focused on ...

30/06/2026

June 29, 2026

Scripps Research scientists demonstrate a faster, cheaper route to making critical drugs using common table sugar New method illustrates how to build a tough ch...

29/06/2026

Op-Ed: Why the 2026 World Cup Is Redefining the Economics of Live Sports Production

By Andy Rayner, CTO, Appear The 2026 FIFA World Cup is the largest football tou...

29/06/2026

Study: Esports Plays Major Role in Gen Z Media Habits, Purchasing Behavior

A new multi-country study from ESL FACEIT Group, Hero Esports, and Niko Partners estimates that 400 million Gen Z consumers regularly engage with esports, under...

29/06/2026

ESPN Sets Multiplatform Plans for America 250 Celebration

ESPN will mark America's 250th anniversary with a series of content initiatives across its linear, digital, and streaming platforms, including a special edi...

29/06/2026

OBSBOT Named Official Camera and Webcam Partner of Esports World Cup 2026

The Esports Foundation has named OBSBOT the Official Camera and Webcam Partner for the Esports World Cup 2026, bringing the company's AI-powered imaging tec...

29/06/2026

Insight Productions Launches Insight Storm, 53-Foot Esports Broadcast Truck

Insight Productions has launched Insight Storm, a 53-foot mobile broadcast unit designed specifically for esports production, live entertainment, and digital-fi...

29/06/2026

Gravity Media Delivers Global Distribution, Streaming Services for World Economic Forum in Dalian

Gravity Media once again provided broadcast, streaming, and content-distribution...

29/06/2026

Wimbledon Introduces AI-Powered Fan Features, Modernized Digital Platforms for 2026 Championships

The All England Lawn Tennis Club and IBM have introduced new and enhanced digita...

29/06/2026

Evolve Dark Matter from Excite Audio

Four-layer instrument aimed at dark electronic music Excite Audio's latest software instrument has been designed with dark drum and bass, atmospheric te...

29/06/2026

Tracktion unleashes Waveform 14 DAW

New AI Assistant, Multi-channel Audio, ARA2 improvements & more Tracktion's DAW software has just received its latest major update, gaining a selection ...

29/06/2026

Focusrite publish 2026 Sustainability Report

Details environmental policies & results The Focusrite Group have just announced that following a long audit process, they have published their 2026 sustain...

29/06/2026

Comcast to Spin Off NBCUniversal, Sky

Share Copy link Facebook X Linkedin Bluesky Email...