Sony Pixel Power calrec Sony

Decoding How NVIDIA AI Workbench Powers App Development

19/06/2024

Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible and showcases new hardware, software, tools and accelerations for NVIDIA RTX PC and workstation users.

The demand for tools to simplify and optimize generative AI development is skyrocketing. Applications based on retrieval-augmented generation (RAG) - a technique for enhancing the accuracy and reliability of generative AI models with facts fetched from specified external sources - and customized models are enabling developers to tune AI models to their specific needs.

While such work may have required a complex setup in the past, new tools are making it easier than ever.

NVIDIA AI Workbench simplifies AI developer workflows by helping users build their own RAG projects, customize models and more. It's part of the RTX AI Toolkit - a suite of tools and software development kits for customizing, optimizing and deploying AI capabilities - launched at COMPUTEX earlier this month. AI Workbench removes the complexity of technical tasks that can derail experts and halt beginners.

What Is NVIDIA AI Workbench? Available for free, NVIDIA AI Workbench enables users to develop, experiment with, test and prototype AI applications across GPU systems of their choice - from laptops and workstations to data center and cloud. It offers a new approach for creating, using and sharing GPU-enabled development environments across people and systems.

A simple installation gets users up and running with AI Workbench on a local or remote machine in just minutes. Users can then start a new project or replicate one from the examples on GitHub. Everything works through GitHub or GitLab, so users can easily collaborate and distribute work. Learn more about getting started with AI Workbench.

How AI Workbench Helps Address AI Project Challenges Developing AI workloads can require manual, often complex processes, right from the start.

Setting up GPUs, updating drivers and managing versioning incompatibilities can be cumbersome. Reproducing projects across different systems can require replicating manual processes over and over. Inconsistencies when replicating projects, like issues with data fragmentation and version control, can hinder collaboration. Varied setup processes, moving credentials and secrets, and changes in the environment, data, models and file locations can all limit the portability of projects.

AI Workbench makes it easier for data scientists and developers to manage their work and collaborate across heterogeneous platforms. It integrates and automates various aspects of the development process, offering:

Ease of setup: AI Workbench streamlines the process of setting up a developer environment that's GPU-accelerated, even for users with limited technical knowledge.

Seamless collaboration: AI Workbench integrates with version-control and project-management tools like GitHub and GitLab, reducing friction when collaborating.

Consistency when scaling from local to cloud: AI Workbench ensures consistency across multiple environments, supporting scaling up or down from local workstations or PCs to data centers or the cloud.

RAG for Documents, Easier Than Ever NVIDIA offers sample development Workbench Projects to help users get started with AI Workbench. The hybrid RAG Workbench Project is one example: It runs a custom, text-based RAG web application with a user's documents on their local workstation, PC or remote system.

Every Workbench Project runs in a container - software that includes all the necessary components to run the AI application. The hybrid RAG sample pairs a Gradio chat interface frontend on the host machine with a containerized RAG server - the backend that services a user's request and routes queries to and from the vector database and the selected large language model.

This Workbench Project supports a wide variety of LLMs available on NVIDIA's GitHub page. Plus, the hybrid nature of the project lets users select where to run inference.

Workbench Projects let users version the development environment and code. Developers can run the embedding model on the host machine and run inference locally on a Hugging Face Text Generation Inference server, on target cloud resources using NVIDIA inference endpoints like the NVIDIA API catalog, or with self-hosting microservices such as NVIDIA NIM or third-party services.

The hybrid RAG Workbench Project also includes:

Performance metrics: Users can evaluate how RAG- and non-RAG-based user queries perform across each inference mode. Tracked metrics include Retrieval Time, Time to First Token (TTFT) and Token Velocity.

Retrieval transparency: A panel shows the exact snippets of text - retrieved from the most contextually relevant content in the vector database - that are being fed into the LLM and improving the response's relevance to a user's query.

Response customization: Responses can be tweaked with a variety of parameters, such as maximum tokens to generate, temperature and frequency penalty.

To get started with this project, simply install AI Workbench on a local system. The hybrid RAG Workbench Project can be brought from GitHub into the user's account and duplicated to the local system.

More resources are available in the AI Decoded user guide. In addition, community members provide helpful video tutorials, like the one from Joe Freeman below.

Customize, Optimize, Deploy Developers often seek to customize AI models for specific use cases. Fine-tuning, a technique that changes the model by training it with additional data, can be useful for style transfer or changing model behavior. AI Workbench helps with fine-tuning, as well.

The Llama-factory AI Workbench Project enables QLoRa, a fine-tuning method that minimizes memory requirements, for a variety of models, as well as
LINK: https://blogs.nvidia.com/blog/ai-decoded-workbench-hybrid-rag/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

14/05/2026

Code of Silence wins Best Drama Series at the 2026 BAFTA's!

Code of Silence has won the BAFTA for Best Drama Series at Sunday night's ceremony at the Royal Festival Hall. The series, starring Rose Ayling-Ellis and w...

14/05/2026

The Wraith Shield Advantage: Transforming L3Harris Radios into AI-Enabled Counter-UAS Sensors

Soldiers equipped with Falcon IV radios will soon gain a sense-and-protect capa...

14/05/2026

Getting into the Space Nuclear Power Game with Next-Generation Technology

Artists concept of the L3Harris Next Gen RTG in flight configuration, designed to provide 250 watts of reliable power for decades-long missions in deep space....

14/05/2026

Vivid Broadcast builds remote production network around Calrec

Vivid Broadcast was embracing remote production long before it became the industry norm. Now, with Calrec's True Control 2.0-enabled Argo M and Type R conso...

14/05/2026

Nielsen data shows NZ vehicle advertisers are shifting gears as fuel pressures make EVs and hybrids an increasingly attractive option

Car ad spend rises sharply in March as more auto buyers turn to electric, hybrid...

14/05/2026

Is This the Year for Agentic AI's Breakout in Broadcast?

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

CBS LA, Los Angeles Rams Ink New TV Deal

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

CueScripts CueiT 4 0 Wins Futures Best of Show Award Pres...

CueScript's CueiT 4.0 Wins Future's Best of Show Award, Presented at 2026 NAB Show by TV Tech CueScript, a leading international developer of professio...

14/05/2026

Expert-Led Education Sessions and Development of Online T...

Expert-Led Education Sessions and Development of Online Training Program Accelerate IPMX Adoption and Deployment The Alliance for IP Media Solutions (AIMS) to...

14/05/2026

Klvr rechargeable battery launches in the USA to cut cost...

Klvr is launching in the United States with a professional-grade rechargeable battery solution that cuts costs and improves performance across live entertainmen...

14/05/2026

Shooting into the depths of Bedlam with URSA Cine 17K 65

Shooting into the depths of Bedlam with URSA Cine 17K 65 Brie Clayton May 14, 2026 0 Comments Indie feature film paired digital 65mm capture with a Bl...

14/05/2026

WeMakeColor expands with Baselight, becoming hybrid color facility

WeMakeColor expands with Baselight, becoming hybrid color facility Caroline Shawley May 14, 2026 0 Comments Boutique Mexican-based studio integrates B...

14/05/2026

Berklee's Summer in the City Returns with Free Concerts Throughout Boston Area

Berklee's Summer in the City Returns with Free Concerts Throughout Boston Ar...

14/05/2026

Chelsey Green Named to Billboard's 2026 Women in Music List

Chelsey Green Named to Billboard's 2026 Women in Music List The Berklee professor and chair of the Recording Academy Board of Trustees joins other high-pr...

14/05/2026

Parks: Tubi, Roku Channel Are Top U.S. FAST Platforms

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

Study: Downstream Fiber Usage Outpaces Cable Broadband

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

NABLF to Honor Kidde With Corporate Leadership Award

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

LABF Awards Four 2026 Preservation Grants

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

Netflix Expands NFL Deal to Five Games

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

Scripps Seals Local Broadcast Deal with Detroit Pistons

Share Copy link Facebook X Linkedin Bluesky Email...

14/05/2026

Glensound marks 60 years of audio innovation at Broadcast...

Six decades of products built around the people who use them...

14/05/2026

Vivid Broadcast builds agile remote production network ar...

Designed to embrace multiple production processes and deliver high-end live broadcast workflows, Vivid Broadcast has combined Calrec True Control 2.0 enabled co...

14/05/2026

Jigsaw24 and EVS Partner to Strengthen Future Proof UK Li...

Today, Jigsaw24, the UK's leading media equipment supplier and systems integrator, announces a new partnership with EVS, a global leader in live video techn...

14/05/2026

Transforming bold ideas into market-ready productions: Digital Originals returns

Transforming bold ideas into market-ready productions: Digital Originals returns 14 May 2026 Top (L-R): Bronte Gosper (Musket), Mema Munro (Rogue One), Nathan ...

14/05/2026

Berklee Convenes Leaders in AI, Music for Inaugural AIMS Symposium

Berklee Convenes Leaders in AI, Music for Inaugural AIMS Symposium The three-day event puts musicians at the center of the future of music creation, ethics, a...

14/05/2026

The Last Laugh: NIAJ Fest Wraps Another Successful Los Angeles Takeover

Back to All News The Last Laugh: NIAJ Fest Wraps Another Successful Los Angeles Takeover Entertainment 14 May 2026 GlobalUnited States Link copied to clipb...

14/05/2026

May 13, 2026

Scripps Research establishes endowed chair honoring renowned structural biologist Ian Wilson Keren Lasker to be inaugural chair holder May 13, 2026 LA JOLLA ...

14/05/2026

Discover hidden havens for city wildlife as Back From The Brink returns for a fifth season

Back From The Brink airs Sunday 17 May and Sunday 24 May at 6.30pm on RT One an...

14/05/2026

Sea You in the Cloud: Subnautica 2' Early Access Dives Onto GeForce NOW

Dive masks on - Subnautica 2 is making a splash on GeForce NOW day-and-date with launch, so members can plunge into the title's brand-new alien ocean from a...

13/05/2026

New Adobe Premiere Color Grading Mode Accelerated on NVIDIA GPUs

New Adobe Premiere Color Grading Mode Accelerated on NVIDIA GPUs Joel Pennington May 13, 2026 0 Comments New NVIDIA RTX-accelerated features streamlin...

13/05/2026

dB Broadcast Delivers New IP-based Cloudbass Sports OB Tr...

Grass Valley announced that dB Broadcast has delivered new IP-based outside broadcast (OB) trucks for Cloudbass, featuring Grass Valley LDX 100 Series cameras a...

13/05/2026

Ikegami Announces its Broadcast Asia 2026 Innovations

Ikegami will exhibit the latest additions to its wide range of broadcast production cameras, control units, viewfinders and monitors on stand 5D3-1 at Broadcast...

13/05/2026

XRSA and FISE Partner to Deliver Immersive Action Sports...

FISE, working with the founding members of the XR Sports Alliance (XRSA), Accedo, Qualcomm Technologies, Inc. and HBS, have collaborated to develop an immersive...

13/05/2026

Canon Unveils New EOS R6 V Full-Frame EOS Camera and RF20-50mm F4 L IS USM PZ Built-In Power Zoom Lens

Canon Unveils New EOS R6 V Full-Frame EOS Camera and RF20-50mm F4 L IS USM PZ Bu...

13/05/2026

Boston Conservatory at Berklee Honors Beth Morrison and Moses Pendleton at Commencement Ceremony

Boston Conservatory at Berklee Honors Beth Morrison and Moses Pendleton at Comme...

13/05/2026

CBS Boston to Air CCBL Baseball

Share Copy link Facebook X Linkedin Bluesky Email...

13/05/2026

Study: Free Streaming Emerges as TV's New Normal

Share Copy link Facebook X Linkedin Bluesky Email...

13/05/2026

Amagi Announces Major Enhancements To CLOUDPORT

Share Copy link Facebook X Linkedin Bluesky Email...

13/05/2026

FCC Releases Updates to TVStudy Software

Share Copy link Facebook X Linkedin Bluesky Email...

13/05/2026

Nexstar Names Elizabeth Ryder EVP, General Counsel

Share Copy link Facebook X Linkedin Bluesky Email...

13/05/2026

Foundry releases Nuke Stage- simplifying virtual producti...

Creative software developer Foundry today announced the latest developments on Nuke Stage. A purpose-built application for end-to-end virtual production and in-...

13/05/2026

Netflix Announces Danish Adaptation of Love is Blind'

Back to All News Netflix Announces Danish Adaptation of Love is Blind' Entertainment 13 May 2026 GlobalDenmarkSweden Link copied to clipboard Love is...

13/05/2026

The Trailer for the Second Season of 'My Family' Only on Netflix June 10

Back to All News The Trailer for the Second Season of My Family Only on Netflix June 10 Entertainment 13 May 2026 GlobalItaly Link copied to clipboard The...

13/05/2026

Netflix Upfront 2026: Get Closer

Back to All News Netflix Upfront 2026: Get Closer Business 13 May 2026 GlobalUnited States Link copied to clipboard Download all assets At our fourth Upf...

13/05/2026

Inter Venezuela Taps Harmonic for PON-Based Mobile Backhaul Service to Support 5G Growth

SAN JOSE, Calif. - May 13, 2026 - Harmonic (NASDAQ: HLIT) today announced that I...

13/05/2026

Tradfluencer - The Sharon Shannon Story

A definitive portrait of one of Ireland's most influential musicians New TV documentary airs Monday 18 May on RT One and RT Player at 9.35pm Watch the...