
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible and showcases new hardware, software, tools and accelerations for NVIDIA RTX PC and workstation users.
The demand for tools to simplify and optimize generative AI development is skyrocketing. Applications based on retrieval-augmented generation (RAG) - a technique for enhancing the accuracy and reliability of generative AI models with facts fetched from specified external sources - and customized models are enabling developers to tune AI models to their specific needs.
While such work may have required a complex setup in the past, new tools are making it easier than ever.
NVIDIA AI Workbench simplifies AI developer workflows by helping users build their own RAG projects, customize models and more. It's part of the RTX AI Toolkit - a suite of tools and software development kits for customizing, optimizing and deploying AI capabilities - launched at COMPUTEX earlier this month. AI Workbench removes the complexity of technical tasks that can derail experts and halt beginners.
What Is NVIDIA AI Workbench? Available for free, NVIDIA AI Workbench enables users to develop, experiment with, test and prototype AI applications across GPU systems of their choice - from laptops and workstations to data center and cloud. It offers a new approach for creating, using and sharing GPU-enabled development environments across people and systems.
A simple installation gets users up and running with AI Workbench on a local or remote machine in just minutes. Users can then start a new project or replicate one from the examples on GitHub. Everything works through GitHub or GitLab, so users can easily collaborate and distribute work. Learn more about getting started with AI Workbench.
How AI Workbench Helps Address AI Project Challenges Developing AI workloads can require manual, often complex processes, right from the start.
Setting up GPUs, updating drivers and managing versioning incompatibilities can be cumbersome. Reproducing projects across different systems can require replicating manual processes over and over. Inconsistencies when replicating projects, like issues with data fragmentation and version control, can hinder collaboration. Varied setup processes, moving credentials and secrets, and changes in the environment, data, models and file locations can all limit the portability of projects.
AI Workbench makes it easier for data scientists and developers to manage their work and collaborate across heterogeneous platforms. It integrates and automates various aspects of the development process, offering:
Ease of setup: AI Workbench streamlines the process of setting up a developer environment that's GPU-accelerated, even for users with limited technical knowledge.
Seamless collaboration: AI Workbench integrates with version-control and project-management tools like GitHub and GitLab, reducing friction when collaborating.
Consistency when scaling from local to cloud: AI Workbench ensures consistency across multiple environments, supporting scaling up or down from local workstations or PCs to data centers or the cloud.
RAG for Documents, Easier Than Ever NVIDIA offers sample development Workbench Projects to help users get started with AI Workbench. The hybrid RAG Workbench Project is one example: It runs a custom, text-based RAG web application with a user's documents on their local workstation, PC or remote system.
Every Workbench Project runs in a container - software that includes all the necessary components to run the AI application. The hybrid RAG sample pairs a Gradio chat interface frontend on the host machine with a containerized RAG server - the backend that services a user's request and routes queries to and from the vector database and the selected large language model.
This Workbench Project supports a wide variety of LLMs available on NVIDIA's GitHub page. Plus, the hybrid nature of the project lets users select where to run inference.
Workbench Projects let users version the development environment and code. Developers can run the embedding model on the host machine and run inference locally on a Hugging Face Text Generation Inference server, on target cloud resources using NVIDIA inference endpoints like the NVIDIA API catalog, or with self-hosting microservices such as NVIDIA NIM or third-party services.
The hybrid RAG Workbench Project also includes:
Performance metrics: Users can evaluate how RAG- and non-RAG-based user queries perform across each inference mode. Tracked metrics include Retrieval Time, Time to First Token (TTFT) and Token Velocity.
Retrieval transparency: A panel shows the exact snippets of text - retrieved from the most contextually relevant content in the vector database - that are being fed into the LLM and improving the response's relevance to a user's query.
Response customization: Responses can be tweaked with a variety of parameters, such as maximum tokens to generate, temperature and frequency penalty.
To get started with this project, simply install AI Workbench on a local system. The hybrid RAG Workbench Project can be brought from GitHub into the user's account and duplicated to the local system.
More resources are available in the AI Decoded user guide. In addition, community members provide helpful video tutorials, like the one from Joe Freeman below.
Customize, Optimize, Deploy Developers often seek to customize AI models for specific use cases. Fine-tuning, a technique that changes the model by training it with additional data, can be useful for style transfer or changing model behavior. AI Workbench helps with fine-tuning, as well.
The Llama-factory AI Workbench Project enables QLoRa, a fine-tuning method that minimizes memory requirements, for a variety of models, as well as
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
03/04/2026
Michigan's Fab Five will reunite for an alternate presentation of the Mich...
03/04/2026
Avid will exhibit at NAB Show 2026 (April 18-22, Booth N2226, Las Vegas Convention Center), demonstrating its Content Core platform and new AI-driven workflow c...
03/04/2026
Mark Roberts Motion Control (MRMC) has announced the appointment of Nick Barthee as Chief Operating Officer.
The announcement follows MRMC's transition fro...
03/04/2026
Interra Systems has announced that Elite Media Technologies has selected its BATON file-based QC solution for media workflows. Elite Media Technologies speciali...
03/04/2026
Ateme has announced that Moldtelecom has deployed Ateme technologies across its streaming workflow, covering encoding, delivery, operations, and analytics.
Mol...
03/04/2026
Grass Valley will demonstrate Framelight X, its content management platform, at NAB Show 2026. The platform connects capture, ingest, editing, and publishing in...
03/04/2026
Encompass Digital Media and Techex have announced a cloud-native Master Control ...
03/04/2026
Live Vertical Video automatically track the action on the court via AI technology and delivers a fully optimized, 9 16 live feed for viewers...
03/04/2026
As the Illini make their first trip to college basketball's biggest stage si...
03/04/2026
After last summer's Softball National Championship victory and last week'...
03/04/2026
The University of Arizona's Men's Basketball team has only loss twice th...
03/04/2026
Eight games across four tournaments will be played in three venues; accommodatio...
03/04/2026
The Ottawa Senators and Bell Media have announced a long-term rights extension for regional Ottawa Senators games on TSN and RDS. TSN Radio 1200 remains the exc...
03/04/2026
Massive production in Phoenix running out of Flagship Mobile unit, Features 50+ ...
03/04/2026
Iconic guitar pedals now available in plug-in form
Guitar effects experts Electro-Harmonix have teamed up with MixWave to turn a collection of their most pr...
03/04/2026
New multi-band AUv3 plug-in announced
Fred Anton Corvest (FAC) offer an extensive range of AUv3 plug-ins and iOS/iPadOS Apps, and their multiband effects pr...
03/04/2026
Just 84 units to be released in the US
Experimental synthesizer and sound-machine extraordinaires SOMA Laboratory have revealed an upcoming special-edition ...
03/04/2026
Emulates the input section of an Ampex 350
One of the latest arrivals to the Iconic Instruments range delivers a new tube preamp plug-in inspired by the cir...
03/04/2026
New York April 2, 2026 TelevisaUnivision, the world's leading Spanish-la...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
03/04/2026
CVP, one of Europe's leading suppliers of professional video and broadcast solutions, today announces the launch of its new German operation and the formati...
03/04/2026
Mark Roberts Motion Control (MRMC) today announces the appointment of Nick Barthee as Chief Operating Officer, strengthening its leadership as the company conti...
03/04/2026
Net Insight introduces programmable Trust Boundaries that make live media interconnection predictable as traffic moves between facilities, networks and cloud en...
03/04/2026
Winning in the new media economy: Avid showcases AI-powered, connected intellige...
03/04/2026
NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibilit...
03/04/2026
NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances
Brie Clayton April 2, 2026
0 Comments
PlayBox Neo will showcase the latest i...
03/04/2026
For Taku Hirano, Everything Is Connected From touring and composition to teaching and instrument design, the in-demand percussionist sees it all as one body o...
03/04/2026
Berklee Honors Humberto Ramirez with Master of Latin Music Award The alumnus and acclaimed trumpeter is honored for his influence as a performer, composer, an...
03/04/2026
VIZ Media Lands Rumiko Takahashi's MAO, Sets April 4 Premiere on Hulu in the...
03/04/2026
Back to All News
Competition Heats Up with Intrigue and Spices: Netflix Unveils...
03/04/2026
Back to All News
Radioactive Emergency Ranks #1 On Netflix's Global Top 10 ...
02/04/2026
HBO and NFL Films have announced Hard Knocks: Training Camp with the Seattle Sea...
02/04/2026
Haivision has announced the Makito ONE, a single-blade video encoding and decoding platform, at NAB Show 2026. The platform combines dual-channel video encoding...
02/04/2026
Telestream has introduced UP.Lens, a cloud-based multiviewer and monitoring serv...
02/04/2026
Mark Roberts Motion Control (MRMC) will exhibit at NAB Show 2026 (Booth C5220, April 19-22, Las Vegas Convention Center), marking the company's 60th anniver...
02/04/2026
Net Insight has introduced programmable Trust Boundaries, a feature integrated i...
02/04/2026
Bitmovin has announced support for SGAI (Server-Guided Ad Insertion) in its playback products, using HLS interstitials. SGAI combines elements of client-side an...
02/04/2026
Riedel Communications' SimplyLive RiMotion R12 replay system is supporting B...
02/04/2026
LTN, a managed IP video transport company, and Ateme, a video compression and de...
02/04/2026
Harmonic has announced that TDF, a broadcast infrastructure operator in France, has deployed Harmonic's XOS Advanced Media Processor and ProStream X Video S...
02/04/2026
Eluvio and the United Rugby Championship (URC) have announced first-year results...