
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, software, tools and accelerations for GeForce RTX PC and NVIDIA RTX workstation users.
Large language models (LLMs) are reshaping productivity. They're capable of drafting documents, summarizing web pages and, having been trained on vast quantities of data, accurately answering questions about nearly any topic.
LLMs are at the core of many emerging use cases in generative AI, including digital assistants, conversational avatars and customer service agents.
Many of the latest LLMs can run locally on PCs or workstations. This is useful for a variety of reasons: users can keep conversations and content private on-device, use AI without the internet, or simply take advantage of the powerful NVIDIA GeForce RTX GPUs in their system. Other models, because of their size and complexity, do no't fit into the local GPU's video memory (VRAM) and require hardware in large data centers.
However, Iit i's possible to accelerate part of a prompt on a data-center-class model locally on RTX-powered PCs using a technique called GPU offloading. This allows users to benefit from GPU acceleration without being as limited by GPU memory constraints.
Size and Quality vs. Performance There's a tradeoff between the model size and the quality of responses and the performance. In general, larger models deliver higher-quality responses, but run more slowly. With smaller models, performance goes up while quality goes down.
This tradeoff isn't always straightforward. There are cases where performance might be more important than quality. Some users may prioritize accuracy for use cases like content generation, since it can run in the background. A conversational assistant, meanwhile, needs to be fast while also providing accurate responses.
The most accurate LLMs, designed to run in the data center, are tens of gigabytes in size, and may not fit in a GPU's memory. This would traditionally prevent the application from taking advantage of GPU acceleration.
However, GPU offloading uses part of the LLM on the GPU and part on the CPU. This allows users to take maximum advantage of GPU acceleration regardless of model size.
Optimize AI Acceleration With GPU Offloading and LM Studio LM Studio is an application that lets users download and host LLMs on their desktop or laptop computer, with an easy-to-use interface that allows for extensive customization in how those models operate. LM Studio is built on top of llama.cpp, so it's fully optimized for use with GeForce RTX and NVIDIA RTX GPUs.
LM Studio and GPU offloading takes advantage of GPU acceleration to boost the performance of a locally hosted LLM, even if the model can't be fully loaded into VRAM.
With GPU offloading, LM Studio divides the model into smaller chunks, or subgraphs, which represent layers of the model architecture. Subgraphs aren't permanently fixed on the GPU, but loaded and unloaded as needed. With LM Studio's GPU offloading slider, users can decide how many of these layers are processed by the GPU.
LM Studio's interface makes it easy to decide how much of an LLM should be loaded to the GPU. For example, imagine using this GPU offloading technique with a large model like Gemma 2 27B. 27B refers to the number of parameters in the model, informing an estimate as to how much memory is required to run the model.
According to 4-bit quantization, a technique for reducing the size of an LLM without significantly reducing accuracy, each parameter takes up a half byte of memory. This means that the model should require about 13.5 billion bytes, or 13.5GB - plus some overhead, which generally ranges from 1-5GB.
Accelerating this model entirely on the GPU requires 19GB of VRAM, available on the GeForce RTX 4090 desktop GPU. With GPU offloading, the model can run on a system with a lower-end GPU and still benefit from acceleration.
The table above shows how to run several popular models of increasing size across a range of GeForce RTX and NVIDIA RTX GPUs. The maximum level of GPU offload is indicated for each combination. Note that even with GPU offloading, users still need enough system RAM to fit the whole model. In LM Studio, it's possible to assess the performance impact of different levels of GPU offloading, compared with CPU only. The below table shows the results of running the same query across different offloading levels on a GeForce RTX 4090 desktop GPU.
Depending on the percent of the model offloaded to GPU, users see increasing throughput performance compared with running on CPUs alone. For the Gemma 2 27B model, performance goes from an anemic 2.1 tokens per second to increasingly usable speeds the more the GPU is used. This enables users to benefit from the performance of larger models that they otherwise would've been unable to run. On this particular model, even users with an 8GB GPU can enjoy a meaningful speedup versus running only on CPUs. Of course, an 8GB GPU can always run a smaller model that fits entirely in GPU memory and get full GPU acceleration.
Achieving Optimal Balance LM Studio's GPU offloading feature is a powerful tool for unlocking the full potential of LLMs designed for the data center, like Gemma 2 27B, locally on RTX AI PCs. It makes larger, more complex models accessible across the entire lineup of PCs powered by GeForce RTX and NVIDIA RTX GPUs.
Download LM Studio to try GPU offloading on larger models, or experiment with a variety of RTX-accelerated LLMs running locally on RTX AI PCs and workstations.
Generative AI is transforming gaming, videoconferencing and interactive experiences of all kinds. Make sense of what's new and what's next by subscribing to the AI Decoded newsletter.
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
09/10/2026
September 10 2026, 06:00 (PDT) Dolby Expands Dolby OptiView Platform with New Capabilities at IBC 2026
New Sports Intelligence helps providers better unders...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
28/09/2026
The NHL has announced Adobe as an Official Creative, Marketing and AI Partner beginning with the 2026-27 season. The partnership will provide the league and its...
28/09/2026
MultiDyne will exhibit at CABSAT 2026 (October 5-7, Dubai Exhibition Centre), showcasing an expanded technology portfolio following its June 2026 acquisition of...
28/09/2026
ARRI and HONOR have announced the HONOR Magic9 Pro Max, a flagship smartphone carrying ARRI branding on the device and packaging as the first product-level expr...
28/09/2026
Net Insight has added the Nimbra 1020G to its pan-Asian live media trial network, giving customers and partners in the region a way to evaluate ST 2110 and JPEG...
28/09/2026
Vindral has been named a recipient of a 2026 Technology and Engineering Emmy Award by the National Academy of Television Arts and Sciences in the category of We...
28/09/2026
Interra Systems will exhibit at CABSAT 2026 (Stand B2-37, Dubai Exhibition Centre, October 5-7), demonstrating its BATON QC platform, ORION monitoring solutions...
28/09/2026
Two years into its new production facility, the TOUR is expanding centralized wo...
28/09/2026
In addition to rolling out a new mobile unit for ESPN's Monday Night Countd...
28/09/2026
Bleacher Report and WWE have announced a multi-year content partnership giving Bleacher Report global highlight rights across WWE's programming slate, inclu...
28/09/2026
Behind The Mic provides a roundup of recent news regarding on-air talent, includ...
28/09/2026
FuboTV has announced a multi-year distribution agreement with the NHL to carry four regional sports networks covering the Carolina Hurricanes, Columbus Blue Jac...
28/09/2026
Program Productions (PPI) has announced the appointment of Alan Ostfield as President and Chief Executive Officer, effective September 28, 2026. Ostfield succee...
28/09/2026
Director Josef Kubota Wladyka attends the premiere of Ha-Chan, Shake Your Booty! with the film's cast and crew at Eccles Theatre on January 22, 2026, in P...
28/09/2026
EQ plug-in can now be played via MIDI
Scaler Music have recently released a new and improved version of their patented musical EQ plug-in, kitting it out wi...
28/09/2026
New synth developed in collaboration with philterSoup
Boutique effects pedal and Eurorack manufacturer Animal Factory Amps have announced that the new compa...
28/09/2026
Rohde & Schwarz adds an integrated multi-channel pulse analysis option to its FS...
28/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
28/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
28/09/2026
Autoscript, the leading global provider of professional teleprompting solutions, today announced Voice Controller for WinPlus-IP, a new voice-controlled prompti...
28/09/2026
IBC 2026 Technical Conference Highlights
David Kirk September 28, 2026
0 Comments
Hero image courtesy Deposit Photos
David Kirk reports from the 11-1...
28/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
28/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
28/09/2026
Nakadi, a global streaming technology innovator dedicated to engineering a new era of digital media delivery, today announced its Live to Theater service based ...
28/09/2026
Monday 28 September 2026
Sky launches Full Fibre 1.6 Gbps for millions of homes
Sky has launched a new Full Fibre 1.6 Gbps product, rolling out from today, gi...
28/09/2026
Arvato Systems Highlights It's Role for Automotive Manufacturers and Suppliers
TISAX
Arvato Systems has successfully completed the TISAX Level 3 assess...
28/09/2026
Crimecall, presented by Carla O'Brien, returns tonight for a brand new series at 9.35pm on RT One and RT Player.
In tonight's episode, alongside appe...
28/09/2026
CELEBRATING 40 YEARS AT THE HEART OF NEW IRISH WRITING
LIVE FROM POETRY IRELAND, DUBLIN AND ON RT RADIO 1 FROM 7PM
DETAILS: www.rte.ie/writing | FOLLOW: #r...
27/09/2026
New compact pedal packs in 120 effects
The latest multi-effects pedal from Boss is a compact unit that promises to give any pedalboard a massive boost in s...
26/09/2026
With short turnaround between game selections, ESPN's operations team is pla...
26/09/2026
Prime Video will use GCV's Bird and Magic mobile units for its inaugural WNB...
26/09/2026
Studio show moves on-site as network models playoff coverage on its NBA flagship...
26/09/2026
Ratings Roundup is a rundown of recent ratings news and is derived from press re...
26/09/2026
Douglas Keeve attends the 2025 Sundance Film Festival screening of Unzipped at the Egyptian Theatre. (Photo by Andrew H. Walker/Shutterstock for Sundance Film...
26/09/2026
Launched alongside three new expansion packs
In celebration of their 30th anniversary, Native Instruments have released a major update for their Maschine 3 ...
26/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
26/09/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/09/2026
In-venue and creative video staffers at the professional and collegiate level ha...
25/09/2026
Infront has named Ronald den Hollander as CEO of its Host Broadcast Services (HBS) subsidiary, effective 1 October 2026. The appointment follows the previously ...
25/09/2026
The North American Broadcasters Association (NABA) will hold Cyber University 2.0, a full-day cybersecurity summit for the broadcast and media industries, on Tu...
25/09/2026
Alpha Networks and Ateme have delivered a joint video platform for the digital d...
25/09/2026
Encompass Digital Media has announced the migration of multiple Viaplay free-to-air channels - TV3, TV6, TV8, and TV10 - to its Altitude Scheduling platform. Al...
25/09/2026
NAB Show New York 2026 will take place October 21-22 at the Jacob K. Javits Convention Center in New York City. Registration is open at nabshow.com/newyork.
Th...
25/09/2026
NHL Network has announced its studio programming lineup for the 2026-27 regular ...
25/09/2026
Clear-Com has announced that First Baptist Nashville has upgraded its production communications with a Clear-Com intercom system as part of a sanctuary technolo...
25/09/2026
Broadcast Management Group (BMG) has deployed EVS XT-VIA live video production s...
25/09/2026
BBright, a Hexaglobe Group company, has announced that dotCentral has joined as a sales and support partner in the United States.
BBright's portfolio cover...
25/09/2026
MeyerPro delivered audio, video, broadcast, and LED production for the Acumatica Summit 2026 at the Seattle Convention Center's Summit Building, January 25-...