Sony Pixel Power calrec Sony

From RAG to Richness: Startup Uplevels Retrieval-Augmented Generation for Enterprises

29/08/2024

Well before OpenAI upended the technology industry with its release of ChatGPT in the fall of 2022, Douwe Kiela already understood why large language models, on their own, could only offer partial solutions for key enterprise use cases.

The young Dutch CEO of Contextual AI had been deeply influenced by two seminal papers from Google and OpenAI, which together outlined the recipe for creating fast, efficient transformer-based generative AI models and LLMs.

Soon after those papers were published in 2017 and 2018, Kiela and his team of AI researchers at Facebook, where he worked at that time, realized LLMs would face profound data freshness issues.

They knew that when foundation models like LLMs were trained on massive datasets, the training not only imbued the model with a metaphorical brain for reasoning across data. The training data also represented the entirety of a model's knowledge that it could draw on to generate answers to users' questions.

Kiela's team realized that, unless an LLM could access relevant real-time data in an efficient, cost-effective way, even the smartest LLM wouldn't be very useful for many enterprises' needs.

So, in the spring of 2020, Kiela and his team published a seminal paper of their own, which introduced the world to retrieval-augmented generation. RAG, as it's commonly called, is a method for continuously and cost-effectively updating foundation models with new, relevant information, including from a user's own files and from the internet. With RAG, an LLM's knowledge is no longer confined to its training data, which makes models far more accurate, impactful and relevant to enterprise users.

Today, Kiela and Amanpreet Singh, a former colleague at Facebook, are the CEO and CTO of Contextual AI, a Silicon Valley-based startup, which recently closed an $80 million Series A round, which included NVIDIA's investment arm, NVentures. Contextual AI is also a member of NVIDIA Inception, a program designed to nurture startups. With roughly 50 employees, the company says it plans to double in size by the end of the year.

The platform Contextual AI offers is called RAG 2.0. In many ways, it's an advanced, productized version of the RAG architecture Kiela and Singh first described in their 2020 paper.

RAG 2.0 can achieve roughly 10x better parameter accuracy and performance over competing offerings, Kiela says.

That means, for example, that a 70-billion-parameter model that would typically require significant compute resources could instead run on far smaller infrastructure, one built to handle only 7 billion parameters without sacrificing accuracy. This type of optimization opens up edge use cases with smaller computers that can perform at significantly higher-than-expected levels.

When ChatGPT happened, we saw this enormous frustration where everybody recognized the potential of LLMs, but also realized the technology wasn't quite there yet, explained Kiela. We knew that RAG was the solution to many of the problems. And we also knew that we could do much better than what we outlined in the original RAG paper in 2020.

Integrated Retrievers and Language Models Offer Big Performance Gains The key to Contextual AI's solutions is its close integration of its retriever architecture, the R in RAG, with an LLM's architecture, which is the generator, or G, in the term. The way RAG works is that a retriever interprets a user's query, checks various sources to identify relevant documents or data and then brings that information back to an LLM, which reasons across this new information to generate a response.

Since around 2020, RAG has become the dominant approach for enterprises that deploy LLM-powered chatbots. As a result, a vibrant ecosystem of RAG-focused startups has formed.

One of the ways Contextual AI differentiates itself from competitors is by how it refines and improves its retrievers through back propagation, a process of adjusting algorithms - the weights and biases - underlying its neural network architecture.

And, instead of training and adjusting two distinct neural networks, that is, the retriever and the LLM, Contextual AI offers a unified state-of-the-art platform, which aligns the retriever and language model, and then tunes them both through back propagation.

Synchronizing and adjusting weights and biases across distinct neural networks is difficult, but the result, Kiela says, leads to tremendous gains in precision, response quality and optimization. And because the retriever and generator are so closely aligned, the responses they create are grounded in common data, which means their answers are far less likely than other RAG architectures to include made up or hallucinated data, which a model might offer when it doesn't know an answer.

Our approach is technically very challenging, but it leads to much stronger coupling between the retriever and the generator, which makes our system far more accurate and much more efficient, said Kiela.

Tackling Difficult Use Cases With State-of-the-Art Innovations RAG 2.0 is essentially LLM-agnostic, which means it works across different open-source language models, like Mistral or Llama, and can accommodate customers' model preferences. The startup's retrievers were developed using NVIDIA's Megatron LM on a mix of NVIDIA H100 and A100 Tensor Core GPUs hosted in Google Cloud.

One of the significant challenges every RAG solution faces is how to identify the most relevant information to answer a user's query when that information may be stored in a variety of formats, such as text, video or PDF.

Contextual AI overcomes this challenge through a mixture of retrievers approach, which aligns different retrievers' sub-specialties with the different formats data is stored in.

Contextual AI deploys a combination
LINK: https://blogs.nvidia.com/blog/contextual-ai-retrieval-augmented-genera...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

01/04/2026

DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION

January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION Douyin Users Can Now Create And Share Videos With Stun...

12/03/2026

Gray Stresses Importance of DRM for NextGen TV in FCC Sports Probe

Share Copy link Facebook X Linkedin Bluesky Email...

12/03/2026

Nebraska's HuskerVision Deploy Lawo IP Tech for Studio Upgrade

Share Copy link Facebook X Linkedin Bluesky Email...

12/03/2026

Comcast NBCU, Telemundo Station Group Announce $600,000 In Grants

Share Copy link Facebook X Linkedin Bluesky Email...

12/03/2026

EditShare To Highlight Analytical AI Capabilities At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

12/03/2026

FreeWheel Launches AI Agent Infrastructure

Share Copy link Facebook X Linkedin Bluesky Email...

12/03/2026

COW Jobs: Editor de Vdeo - Direct Response, Performance Ads - Brazil, Remote

COW Jobs: Editor de V deo - Direct Response, Performance Ads - Brazil, Remote Brie Clayton March 11, 2026 0 Comments Editor(a) de V deo (Direct Respon...

12/03/2026

Avatar: Fire and Ash Graded with DaVinci Resolve Studio

Avatar: Fire and Ash Graded with DaVinci Resolve Studio Brie Clayton March 11, 2026 0 Comments Colorist delivers premium cinematic color across 2D, 3D...

12/03/2026

Boston Conservatory to Timothe Chalamet: We Care About Ballet and Opera

Boston Conservatory to Timoth e Chalamet: We Care About Ballet and Opera Boston Conservatory at Berklee students and faculty respond to the actors recent comm...

11/03/2026

Calrec To Unlock Hybrid Workflows At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

Matrox Video Enables the Next Era of Software-Defined Med...

Matrox Video will showcase its vision for the future of live production at NAB 2026 in Las Vegas, April 19-22, highlighting how broadcasters and media organizat...

11/03/2026

GlobalM Showcases Distributed Video Gateway Architecture...

Geneva-based technology company, GlobalM SA, is presenting its GMX Distributed Video Gateway, a software-defined IP media transport platform designed to replace...

11/03/2026

Video is King - 2026 Iconik Media Stats Report Finds Vide...

Backlight (booth #N2829), the company behind Iconik and Wildmoka, which power video workflows for large media and entertainment organizations, has released the ...

11/03/2026

QuickLinks Latest StudioEdge Models to Make North America...

QuickLink, a leading provider of award-winning video production and remote guest contribution solutions, presents its latest StudioEdge models at The NAB Show ...

11/03/2026

Telestream Expands Its Cloud Services with the Introducti...

Telestream, a global leader in media workflow technologies, today announced the expansion of Telestream Cloud Services with the introduction of UP, a new cloud-...

11/03/2026

Operative Launches AOS Configuration for Digital-First Mo...

Operative, the preferred advertising management provider for the world's leading media brands, today announced the launch of AOS for digital media, an AI-po...

11/03/2026

Calrec Redefines Broadcast Workflows at NAB 2026

Calrec will be located in Central Hall, on Booth C6907 Choice without compromise The broadcast industry is going through a rapid evolution that s signalling a...

11/03/2026

Worldstream and Cubbit launch sovereign S3 cloud storage...

The new service is hosted and operated entirely in the Netherlands, combining data sovereignty, resilience, scalability, and predictable costs without relying...

11/03/2026

Ease Live powers interactive Premier Padel experiences on...

Ease Live, an Evertz company and leader in interactive graphical overlays, today announced the successful deployment of its platform on Red Bull TV for Premier ...

11/03/2026

Mediagenix Title Management Accelerates Content Monetizat...

Mediagenix, a global leader in smart content solutions to profitably connect the right content to the right audience, is advancing its Semantic Intelligence cap...

11/03/2026

Emergent Launches Fusion- The Interactive Anything Platfo...

Emergent, a leading provider of AI-enhanced media production solutions, today announced the official launch of Fusion, a powerful, no-code application builder d...

11/03/2026

Techex Names Matt McKee as Senior Director of Sales, Americas

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

IAB Tech Lab Announces Content Monetization Protocol for AI LLMs

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

Mondae Hott Joins Kokusai Denki as Northeast Sales Manager

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

Gray Media to Air Cincinnati Reds' Games on WXIX FOX19

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

Shure Audio Solutions Deliver Super Bowl Win

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

UK's First Live Broadcast Using New n40 Private 5G Spectrum

Share Copy link Facebook X Linkedin Bluesky Email...

11/03/2026

Utah Scientific Expands Technology Partner Program With I...

Utah Scientific today announced the expansion of its Technology Partner Program with the addition of Audinate, Bitfocus, and Skaarhoj, three industry leaders wh...

11/03/2026

DigitalGlue Ends the Post Production Tax creativespace In...

DigitalGlue, creator of the creative.space on-premise managed storage platform, today revealed plans to launch creative.space Intelligence (CSI) at NAB 2026 (Bo...

11/03/2026

Maxon and Tencent Cloud Partner to Integrate HY 3D into C...

Maxon, maker of powerful, approachable software solutions for creators working in 2D and 3D design, motion graphics, visual effects, gaming, and more, has annou...

11/03/2026

NUGEN Audio Halo Vision Plug In Serves as Spatial Compass...

Composer and Re-recording Mixer Michael Phillips Keeley has built his career around immersive storytelling. Working from his Dolby Atmos-equipped studio, Sound ...

11/03/2026

YES selects Synamedia Iris to power advanced advertising

Leading video software provider Synamedia today announced that YES, the pay-TV subsidiary of the telco Bezeq (TASE: BEZQ), has selected Synamedia Iris to delive...

11/03/2026

Cost Savings Scalability and Smarter Monetization Viacces...

As media companies face increasing cost pressures and operational complexity, at the 2026 NAB Show in Las Vegas, Viaccess-Orca (VO), a global leader in OTT / TV...

11/03/2026

Digital Alert Systems Unveils Version 6 Software for DASD...

Digital Alert Systems, a global leader in emergency communications solutions for media providers, today announced the release of Version 6 software for its DASD...

11/03/2026

SES Brings Satellite Connectivity to Refugees in Chad

First Medium-Earth Orbit (MEO) deployment of the emergency.lu platform for refugees and their host communities' use provides dependable broadband for humani...

11/03/2026

Foundry releases Nuke 17.0

Foundry releases Nuke 17.0 Brie Clayton March 1, 2026 0 Comments Native Gaussian Splat support, new 3D system based on USD, expanded machine learning ca...

11/03/2026

Preserving UNESCO World Heritage with URSA Cine Immersive

Preserving UNESCO World Heritage with URSA Cine Immersive Brie Clayton March 1, 2026 0 Comments The Explorers turned to France's cultural landmark...

11/03/2026

I Clicked This By Accident And It Made After Effects SO Much Faster

I Clicked This By Accident And It Made After Effects SO Much Faster Graham Quince March 1, 2026 0 Comments Discover how Region of Interest in Adobe A...

11/03/2026

Cine Gear Connect Brings a Focused All-Day Experience to Industry City, NY

Cine Gear Connect Brings a Focused All-Day Experience to Industry City, NY Brie Clayton March 4, 2026 0 Comments Registration is now open for Cine Gea...

11/03/2026

La Vorgine Edited and Finished with DaVinci Resolve Studio

La Vor gine Edited and Finished with DaVinci Resolve Studio Brie Clayton March 4, 2026 0 Comments One of Colombia's most ambitious projects goes g...

11/03/2026

SoundMarket Launches 18,000+ Tracks of Real Music by Award-Winning Composers for Editors and Post Professionals

SoundMarket Launches 18,000 Tracks of Real Music by Award-Winning Composers for...

11/03/2026

Capta Center Supports NOVO19 Remote Production with Blackmagic Design

Capta Center Supports NOVO19 Remote Production with Blackmagic Design Brie Clayton March 5, 2026 0 Comments The facility provides production and playo...

11/03/2026

DigitalGlue Ends the Post-Production Tax: creative.space Intelligence (CSI) Unifies On-Premise Storage with Forensic AI at NAB 2026

DigitalGlue Ends the Post-Production Tax: creative.space Intelligence (CSI) Unif...

11/03/2026

Kochi Sun Sun Uses Blackmagic Replay for High School Volleyball Finals

Kochi Sun Sun Uses Blackmagic Replay for High School Volleyball Finals Brie Clayton March 9, 2026 0 Comments Versatile Blackmagic Replay system proves...

11/03/2026

Richard Bona Joins Berklee for Signature Series Concert

Richard Bona Joins Berklee for Signature Series Concert The Grammy-winning Cameroonian bassist and vocalist collaborates with students and faculty in a progra...

11/03/2026

New NVIDIA Nemotron 3 Super Delivers 5x Higher Throughput for Agentic AI

Launched today, NVIDIA Nemotron 3 Super is a 120 billion parameter open model with 12 billion active parameters designed to run complex agentic AI systems at sc...