Sony Pixel Power calrec Sony

How to Accelerate Larger LLMs Locally on RTX With LM Studio

23/10/2024

Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, software, tools and accelerations for GeForce RTX PC and NVIDIA RTX workstation users.

Large language models (LLMs) are reshaping productivity. They're capable of drafting documents, summarizing web pages and, having been trained on vast quantities of data, accurately answering questions about nearly any topic.

LLMs are at the core of many emerging use cases in generative AI, including digital assistants, conversational avatars and customer service agents.

Many of the latest LLMs can run locally on PCs or workstations. This is useful for a variety of reasons: users can keep conversations and content private on-device, use AI without the internet, or simply take advantage of the powerful NVIDIA GeForce RTX GPUs in their system. Other models, because of their size and complexity, do no't fit into the local GPU's video memory (VRAM) and require hardware in large data centers.

However, Iit i's possible to accelerate part of a prompt on a data-center-class model locally on RTX-powered PCs using a technique called GPU offloading. This allows users to benefit from GPU acceleration without being as limited by GPU memory constraints.

Size and Quality vs. Performance There's a tradeoff between the model size and the quality of responses and the performance. In general, larger models deliver higher-quality responses, but run more slowly. With smaller models, performance goes up while quality goes down.

This tradeoff isn't always straightforward. There are cases where performance might be more important than quality. Some users may prioritize accuracy for use cases like content generation, since it can run in the background. A conversational assistant, meanwhile, needs to be fast while also providing accurate responses.

The most accurate LLMs, designed to run in the data center, are tens of gigabytes in size, and may not fit in a GPU's memory. This would traditionally prevent the application from taking advantage of GPU acceleration.

However, GPU offloading uses part of the LLM on the GPU and part on the CPU. This allows users to take maximum advantage of GPU acceleration regardless of model size.

Optimize AI Acceleration With GPU Offloading and LM Studio LM Studio is an application that lets users download and host LLMs on their desktop or laptop computer, with an easy-to-use interface that allows for extensive customization in how those models operate. LM Studio is built on top of llama.cpp, so it's fully optimized for use with GeForce RTX and NVIDIA RTX GPUs.

LM Studio and GPU offloading takes advantage of GPU acceleration to boost the performance of a locally hosted LLM, even if the model can't be fully loaded into VRAM.

With GPU offloading, LM Studio divides the model into smaller chunks, or subgraphs, which represent layers of the model architecture. Subgraphs aren't permanently fixed on the GPU, but loaded and unloaded as needed. With LM Studio's GPU offloading slider, users can decide how many of these layers are processed by the GPU.

LM Studio's interface makes it easy to decide how much of an LLM should be loaded to the GPU. For example, imagine using this GPU offloading technique with a large model like Gemma 2 27B. 27B refers to the number of parameters in the model, informing an estimate as to how much memory is required to run the model.

According to 4-bit quantization, a technique for reducing the size of an LLM without significantly reducing accuracy, each parameter takes up a half byte of memory. This means that the model should require about 13.5 billion bytes, or 13.5GB - plus some overhead, which generally ranges from 1-5GB.

Accelerating this model entirely on the GPU requires 19GB of VRAM, available on the GeForce RTX 4090 desktop GPU. With GPU offloading, the model can run on a system with a lower-end GPU and still benefit from acceleration.

The table above shows how to run several popular models of increasing size across a range of GeForce RTX and NVIDIA RTX GPUs. The maximum level of GPU offload is indicated for each combination. Note that even with GPU offloading, users still need enough system RAM to fit the whole model. In LM Studio, it's possible to assess the performance impact of different levels of GPU offloading, compared with CPU only. The below table shows the results of running the same query across different offloading levels on a GeForce RTX 4090 desktop GPU.

Depending on the percent of the model offloaded to GPU, users see increasing throughput performance compared with running on CPUs alone. For the Gemma 2 27B model, performance goes from an anemic 2.1 tokens per second to increasingly usable speeds the more the GPU is used. This enables users to benefit from the performance of larger models that they otherwise would've been unable to run. On this particular model, even users with an 8GB GPU can enjoy a meaningful speedup versus running only on CPUs. Of course, an 8GB GPU can always run a smaller model that fits entirely in GPU memory and get full GPU acceleration.

Achieving Optimal Balance LM Studio's GPU offloading feature is a powerful tool for unlocking the full potential of LLMs designed for the data center, like Gemma 2 27B, locally on RTX AI PCs. It makes larger, more complex models accessible across the entire lineup of PCs powered by GeForce RTX and NVIDIA RTX GPUs.

Download LM Studio to try GPU offloading on larger models, or experiment with a variety of RTX-accelerated LLMs running locally on RTX AI PCs and workstations.

Generative AI is transforming gaming, videoconferencing and interactive experiences of all kinds. Make sense of what's new and what's next by subscribing to the AI Decoded newsletter.
LINK: https://blogs.nvidia.com/blog/ai-decoded-lm-studio/...
See more stories from nvidia

More from Nvidia

05/02/2026

GeForce NOW Celebrates Six Years of Streaming With 24 Games in February

Break out the cake and green sprinkles - GeForce NOW is turning six. Since launch, members have streamed over 1 billion hours, and the party's just getting...

04/02/2026

Nemotron Labs: How AI Agents Are Turning Documents Into Real-Time Business Intelligence

Editor's note: This post is part of the Nemotron Labs blog series, which exp...

03/02/2026

Everything Will Be Represented in a Virtual Twin, NVIDIA CEO Jensen Huang Says at 3DEXPERIENCE World

At 3DEXPERIENCE World in Houston, NVIDIA founder and CEO Jensen Huang and Dassau...

29/01/2026

Mercedes-Benz Unveils New S-Class Built on NVIDIA DRIVE AV, Which Enables an L4-Ready Architecture

Mercedes-Benz is marking 140 years of automotive innovation with a new S-Class b...

29/01/2026

Into the Omniverse: Physical AI Open Models and Frameworks Advance Robots and Autonomous Systems

Editor's note: This post is part of Into the Omniverse, a series focused on ...

29/01/2026

GeForce NOW Brings GeForce RTX Gaming to Linux PCs

Get ready to game - the native GeForce NOW app for Linux PCs is now available in beta, letting Linux desktops tap directly into GeForce RTX performance from the...

28/01/2026

Accelerating Science: A Blueprint for a Renewed National Quantum Initiative

Quantum technologies are rapidly emerging as foundational capabilities for economic competitiveness, national security and scientific leadership in the 21st cen...

22/01/2026

NVIDIA DRIVE AV Raises the Bar for Vehicle Safety as Mercedes-Benz CLA Earns Top Euro NCAP Award

AI-powered driver assistance technologies are becoming standard equipment, funda...

22/01/2026

Flight Controls Are Cleared for Takeoff on GeForce NOW

The wait is over, pilots. Flight control support - one of the most community-requested features for GeForce NOW - is live starting today, following its announce...

22/01/2026

From Pilot to Profit: Survey Reveals the Financial Services Industry Is Doubling Down on AI Investment and Open Source

AI has taken center stage in financial services, automating the research and exe...

22/01/2026

How to Get Started With Visual Generative AI on NVIDIA RTX PCs

AI-powered content generation is now embedded in everyday tools like Adobe and Canva, with a slew of agencies and studios incorporating the technology into thei...

21/01/2026

Largest Infrastructure Buildout In Human History': Jensen Huang on AI's Five-Layer Cake' at Davos

From skilled trades to startups, AI's rapid expansion is the beginning of th...

21/01/2026

Largest Infrastructure Buildout In Human History: Jensen Huang on AI's Five-Layer Cake at Davos

From skilled trades to startups, AI's rapid expansion is the beginning of th...

15/01/2026

Survive the Quarantine Zone and More With Devolver Digital Games on GeForce NOW

NVIDIA kicked off the year at CES, where the crowd buzzed about the latest gaming announcements - including the native GeForce NOW app for Linux and Amazon Fire...

13/01/2026

CEOs of NVIDIA and Lilly Share Blueprint for What Is Possible' in AI and Drug Discovery

NVIDIA and Lilly are putting together a blueprint for what is possible in the f...

09/01/2026

NVIDIA Unveils Multi-Agent Intelligent Warehouse and Catalog Enrichment AI Blueprints to Power the Retail Pipeline

Every that was easy shopping moment is made possible by teams working to hit s...

08/01/2026

Japan Science and Technology Agency Develops NVIDIA-Powered Moonshot Robot for Elderly Care

The next universal technology since the smartphone is on the horizon - and it ma...

08/01/2026

AI Copilot Keeps Berkeley's X-Ray Particle Accelerator on Track

In the rolling hills of Berkeley, California, an AI agent is supporting high-stakes physics experiments at the Advanced Light Source (ALS) particle accelerator....

08/01/2026

More Ways to Play, More Games to Love - GeForce NOW Wraps CES With Linux Support, Fire TV App, Flight Stick Controls

NVIDIA is wrapping up a big week at the CES trade show with a set of GeForce NOW...

07/01/2026

From Warehouse to Wallet: New State of AI in Retail and CPG Survey Uncovers How AI Is Rewiring Supply Chains and Customer Experiences

AI has transformed retail and consumer packaged goods (CPG) operations, enhancin...

05/01/2026

NVIDIA Expands Global DRIVE Hyperion Ecosystem to Accelerate the Road to Full Autonomy

At the CES trade show running this week in Las Vegas, NVIDIA announced that the ...

05/01/2026

NVIDIA DGX Spark and DGX Station Power the Latest Open-Source and Frontier Models From the Desktop

Open-source AI is accelerating innovation across industries, and NVIDIA DGX Spar...

05/01/2026

NVIDIA DGX SuperPOD Sets the Stage for Rubin-Based Systems

NVIDIA DGX SuperPOD is paving the way for large-scale system deployments built on the NVIDIA Rubin platform - the next leap forward in AI computing. At the CES...

05/01/2026

NVIDIA BlueField-Powered Cybersecurity and Acceleration Arrive on NVIDIA Enterprise AI Factory Validated Design

AI is powering breakthroughs across industries, helping enterprises operate with...

05/01/2026

NVIDIA Rubin Platform, Open Models, Autonomous Driving: NVIDIA Presents Blueprint for the Future at CES

NVIDIA founder and CEO Jensen Huang took the stage at the Fontainebleau Las Vega...

05/01/2026

NVIDIA DLSS 4.5, Path Tracing and G-SYNC Pulsar Supercharge Gameplay With Enhanced Performance and Visuals

At the CES trade show, NVIDIA today announced DLSS 4.5, which introduces Dynamic...

05/01/2026

NVIDIA RTX Accelerates 4K AI Video Generation on PC With LTX-2 and ComfyUI Upgrades

2025 marked a breakout year for AI development on PC. PC-class small language m...

05/01/2026

NVIDIA Brings GeForce RTX Gaming to More Devices With New GeForce NOW Apps for Linux PC and Amazon Fire TV

Announced at the CES trade show running this week in Las Vegas, NVIDIA is bringi...

01/01/2026

GeForce NOW Rings In 2026 With 14 New Games in January

New year, new games, all with RTX 5080-powered cloud energy. GeForce NOW is kicking off 2026 by looking back at an unforgettable year of wins and wildly high fr...

25/12/2025

Make Spirits Bright With Holiday Hits on GeForce NOW

Holiday lights are twinkling, hot cocoa's on the stove and gamers are settling in for a well-earned break. Whether staying in or heading on a winter getawa...

22/12/2025

Marine Biological Laboratory Explores Human Memory With AI and Virtual Reality

The works of Plato state that when humans have an experience, some level of change occurs in their brain, which is powered by memory - specifically long-term me...

18/12/2025

NVIDIA, US Government to Boost AI Infrastructure and R&D Investments Through Landmark Genesis Mission

NVIDIA will join the U.S. Department of Energy's (DOE) Genesis Mission as a ...

18/12/2025

Now Generally Available, NVIDIA RTX PRO 5000 72GB Blackwell GPU Expands Memory Options for Desktop Agentic AI

Top-notch options for AI at the desktops of developers, engineers and designers ...

18/12/2025

Deck the Vaults: Fallout: New Vegas' Joins the Cloud This Holiday Season

Step out of the vault and into the future of gaming with Fallout: New Vegas streaming on GeForce NOW, just in time to celebrate the newest season of the hit Ama...

17/12/2025

UC San Diego Lab Advances Generative AI Research With NVIDIA DGX B200 System

The Hao AI Lab research team at the University of California San Diego - at the forefront of pioneering AI model innovation - recently received an NVIDIA DGX B...

17/12/2025

Into the Omniverse: OpenUSD and NVIDIA Halos Accelerate Safety for Robotaxis, Physical AI Systems

Editor's note: This post is part of Into the Omniverse, a series focused on ...

15/12/2025

NVIDIA Acquires Open-Source Workload Management Provider SchedMD

NVIDIA today announced it has acquired SchedMD - the leading developer of Slurm, an open-source workload management system for high-performance computing (HPC) ...

15/12/2025

How to Fine-Tune an LLM on NVIDIA GPUs With Unsloth

Modern workflows showcase the endless possibilities of generative and agentic AI on PCs. Of many, some examples include tuning a chatbot to handle product-supp...

12/12/2025

Cheers to AI: ADAM Robot Bartender Makes Drinks at Vegas Golden Knights Game

In Las Vegas's T-Mobile Arena, fans of the Golden Knights are getting more than just hockey - they're getting a taste of the future. ADAM, a robot devel...

11/12/2025

As AI Grows More Complex, Model Builders Rely on NVIDIA

Unveiling what it describes as the most capable model series yet for professional knowledge work, OpenAI launched GPT-5.2 today. The model was trained and deplo...

11/12/2025

Ride Into Adventure With Capcom's Monster Hunter Stories' Series in the Cloud

Hunters, saddle up - adventure awaits in the cloud. Journey into the world of M...

10/12/2025

3 Ways NVIDIA Is Powering the Industrial Revolution

The NVIDIA accelerated computing platform is leading supercomputing benchmarks once dominated by CPUs, enabling AI, science, business and computing efficiency w...

10/12/2025

How NVIDIA H100 GPUs on CoreWeave's AI Cloud Platform Delivered a Record-Breaking Graph500 Run

The world's top-performing system for graph processing at scale was built on...

10/12/2025

Opt-In NVIDIA Software Enables Data Center Fleet Management

As the scale and complexity of AI infrastructure grows, data center operators need continuous visibility into factors including performance, temperature and pow...

04/12/2025

Robots' Holiday Wishes Come True: NVIDIA Jetson Platform Offers High-Performance Edge AI at Festive Prices

Developers, researchers, hobbyists and students can take a byte out of holiday s...

04/12/2025

Game the Halls: GeForce NOW Brings Holiday Cheer With 30 New Games in the Cloud

Editor's note: The Game Pass edition of Hogwarts Legacy' will also be supported on GeForce NOW when the Steam and Epic Games Store versions launch on t...

03/12/2025

Mixture of Experts Powers the Most Intelligent Frontier AI Models, Runs 10x Faster on NVIDIA Blackwell NVL72

The top 10 most intelligent open-source models all use a mixture-of-experts arch...

02/12/2025

NVIDIA Partners With Mistral AI to Accelerate New Family of Open Models

Today, Mistral AI announced the Mistral 3 family of open-source multilingual, multimodal models, optimized across NVIDIA supercomputing and edge platforms. M...

02/12/2025

NVIDIA and AWS Expand Full-Stack Partnership, Providing the Secure, High-Performance Compute Platform Vital for Future Innovation

At AWS re:Invent, NVIDIA and Amazon Web Services expanded their strategic collab...

01/12/2025

At NeurIPS, NVIDIA Advances Open Model Development for Digital and Physical AI

Researchers worldwide rely on open-source technologies as the foundation of their work. To equip the community with the latest advancements in digital and physi...