Sony Pixel Power calrec Sony

TOPS of the Class: Decoding AI Performance on RTX AI PCs and Workstations

12/06/2024

Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, software, tools and accelerations for RTX PC users.

The era of the AI PC is here, and it's powered by NVIDIA RTX and GeForce RTX technologies. With it comes a new way to evaluate performance for AI-accelerated tasks, and a new language that can be daunting to decipher when choosing between the desktops and laptops available.

While PC gamers understand frames per second (FPS) and similar stats, measuring AI performance requires new metrics.

Coming Out on TOPS The first baseline is TOPS, or trillions of operations per second. Trillions is the important word here - the processing numbers behind generative AI tasks are absolutely massive. Think of TOPS as a raw performance metric, similar to an engine's horsepower rating. More is better.

Compare, for example, the recently announced Copilot+ PC lineup by Microsoft, which includes neural processing units (NPUs) able to perform upwards of 40 TOPS. Performing 40 TOPS is sufficient for some light AI-assisted tasks, like asking a local chatbot where yesterday's notes are.

But many generative AI tasks are more demanding. NVIDIA RTX and GeForce RTX GPUs deliver unprecedented performance across all generative tasks - the GeForce RTX 4090 GPU offers more than 1,300 TOPS. This is the kind of horsepower needed to handle AI-assisted digital content creation, AI super resolution in PC gaming, generating images from text or video, querying local large language models (LLMs) and more.

Insert Tokens to Play TOPS is only the beginning of the story. LLM performance is measured in the number of tokens generated by the model.

Tokens are the output of the LLM. A token can be a word in a sentence, or even a smaller fragment like punctuation or whitespace. Performance for AI-accelerated tasks can be measured in tokens per second.

Another important factor is batch size, or the number of inputs processed simultaneously in a single inference pass. As an LLM will sit at the core of many modern AI systems, the ability to handle multiple inputs (e.g. from a single application or across multiple applications) will be a key differentiator. While larger batch sizes improve performance for concurrent inputs, they also require more memory, especially when combined with larger models.

The more you batch, the more (time) you save. RTX GPUs are exceptionally well-suited for LLMs due to their large amounts of dedicated video random access memory (VRAM), Tensor Cores and TensorRT-LLM software.

GeForce RTX GPUs offer up to 24GB of high-speed VRAM, and NVIDIA RTX GPUs up to 48GB, which can handle larger models and enable higher batch sizes. RTX GPUs also take advantage of Tensor Cores - dedicated AI accelerators that dramatically speed up the computationally intensive operations required for deep learning and generative AI models. That maximum performance is easily accessed when an application uses the NVIDIA TensorRT software development kit (SDK), which unlocks the highest-performance generative AI on the more than 100 million Windows PCs and workstations powered by RTX GPUs.

The combination of memory, dedicated AI accelerators and optimized software gives RTX GPUs massive throughput gains, especially as batch sizes increase.

Text-to-Image, Faster Than Ever Measuring image generation speed is another way to evaluate performance. One of the most straightforward ways uses Stable Diffusion, a popular image-based AI model that allows users to easily convert text descriptions into complex visual representations.

With Stable Diffusion, users can quickly create and refine images from text prompts to achieve their desired output. When using an RTX GPU, these results can be generated faster than processing the AI model on a CPU or NPU.

That performance is even higher when using the TensorRT extension for the popular Automatic1111 interface. RTX users can generate images from prompts up to 2x faster with the SDXL Base checkpoint - significantly streamlining Stable Diffusion workflows.

ComfyUI, another popular Stable Diffusion user interface, added TensorRT acceleration last week. RTX users can now generate images from prompts up to 60% faster, and can even convert these images to videos using Stable Video Diffuson up to 70% faster with TensorRT.

TensorRT acceleration can be put to the test in the new UL Procyon AI Image Generation benchmark, which delivers speedups of 50% on a GeForce RTX 4080 SUPER GPU compared with the fastest non-TensorRT implementation.

TensorRT acceleration will soon be released for Stable Diffusion 3 - Stability AI's new, highly anticipated text-to-image model - boosting performance by 50%. Plus, the new TensorRT-Model Optimizer enables accelerating performance even further. This results in a 70% speedup compared with the non-TensorRT implementation, along with a 50% reduction in memory consumption.

Of course, seeing is believing - the true test is in the real-world use case of iterating on an original prompt. Users can refine image generation by tweaking prompts significantly faster on RTX GPUs, taking seconds per iteration compared with minutes on a Macbook Pro M3 Max. Plus, users get both speed and security with everything remaining private when running locally on an RTX-powered PC or workstation.

The Results Are in and Open Sourced But don't just take our word for it. The team of AI researchers and engineers behind the open-source Jan.ai recently integrated TensorRT-LLM into its local chatbot app, then tested these optimizations for themselves.

Source: Jan.ai The researchers tested its implementation of TensorRT-LLM against the open-source llama.cpp inference engine across a variety of GPUs and CPUs used by the community. They found that TensorRT is 30-70% faster than llam
LINK: https://blogs.nvidia.com/blog/ai-decoded-tops/...
See more stories from nvidia

More from Nvidia

17/03/2026

NVIDIA, Telecom Leaders Build AI Grids to Optimize Inference on Distributed Networks

As AI native applications scale to more users, agents and devices, the telecommu...

17/03/2026

Snap Decisions: How Open Libraries for Accelerated Data Processing Boost A/B Testing for Snapchat

The features on social media apps like Snapchat evolve nearly as fast as what...

17/03/2026

GTC Spotlights NVIDIA RTX PCs and DGX Sparks Running Latest Open Models and AI Agents Locally

The paradigm of consumer computing has revolved around the concept of a personal...

12/03/2026

Into the Omniverse: How Industrial AI and Digital Twins Accelerate Design, Engineering and Manufacturing Across Industries

Editor's note: This post is part of Into the Omniverse, a series focused on ...

12/03/2026

GeForce NOW Raises the Game at the Game Developers Conference

GeForce NOW is bringing the game to the Game Developers Conference (GDC), running this week in San Francisco. While developers build the future of gaming, GeFor...

11/03/2026

New NVIDIA Nemotron 3 Super Delivers 5x Higher Throughput for Agentic AI

Launched today, NVIDIA Nemotron 3 Super is a 120 billion parameter open model with 12 billion active parameters designed to run complex agentic AI systems at sc...

10/03/2026

NVIDIA and ComfyUI Streamline Local AI Video Generation for Game Developers and Creators at GDC

Game developers and artists are building cinematic worlds and iconic characters ...

10/03/2026

NVIDIA Virtualizes Game Development With RTX PRO Server

Game development teams are working across larger worlds, more complex pipelines and more distributed teams than ever. At the same time, many studios still rely ...

10/03/2026

As Open Models Spark AI Boom, NVIDIA Jetson Brings It to Life at the Edge

The Cat 306 CR mini-excavator weighs just under eight tons and fits inside a standard shipping container. It's the machine a contractor rents when the job s...

10/03/2026

NVIDIA and Thinking Machines Lab Announce Long-Term Gigawatt-Scale Strategic Partnership

NVIDIA and Thinking Machines Lab announced today a multiyear strategic partnersh...

09/03/2026

How AI Is Driving Revenue, Cutting Costs and Boosting Productivity for Every Industry in 2026

AI is everywhere and accelerating everything - becoming essential infrastructure...

09/03/2026

ABB Robotics Taps NVIDIA Omniverse to Deliver IndustrialGrade Physical AI at Scale

ABB Robotics and NVIDIA today announced a breakthrough partnership that brings i...

05/03/2026

March Into the Cloud With 15 New Games Coming to GeForce NOW

March is in full bloom, and that means a fresh wave of games heading to the cloud. 15 new titles are joining the GeForce NOW library this month. Leading the Ma...

28/02/2026

NVIDIA and Partners Show That Software-Defined AI-RAN Is the Next Wireless Generation

AI-RAN is moving from lab to field, showing that a software-defined approach is ...

28/02/2026

NVIDIA Advances Autonomous Networks With Agentic AI Blueprints and Telco Reasoning Models

Autonomous networks - intelligent, self-managing telecommunications operations -...

26/02/2026

The Nightmare Returns in the Cloud: GeForce NOW Unleashes Capcom's Resident Evil Requiem'

GeForce NOW's anniversary celebration reaches a chilling crescendo as Capcom...

26/02/2026

Horror Awakens in the Cloud: GeForce NOW Unleashes Capcom's Resident Evil: Requiem'

GeForce NOW's anniversary celebration reaches a chilling crescendo as Capcom...

24/02/2026

From Radiology to Drug Discovery, Survey Reveals AI Is Delivering Clear Return on Investment in Healthcare

AI is accelerating every aspect of healthcare - from radiology and drug discover...

23/02/2026

NVIDIA Brings AI-Powered Cybersecurity to World's Critical Infrastructure

As technologies and systems become more digitalized and connected across the world, operational technology (OT) environments and industrial control systems (ICS...

19/02/2026

All About the Games: Play Over 4,500 Titles With GeForce NOW

The GeForce NOW anniversary celebration keeps on rolling, and this week is all about the games that make it possible. With more than 4,500 titles supported in t...

19/02/2026

Survey Reveals AI Advances in Telecom: Networks and Automation in Driver's Seat as Return on Investment Climbs

AI is accelerating the telecommunications industry's transformation, becomin...

17/02/2026

NVIDIA and Global Industrial Software Leaders Partner With India's Largest Manufacturers to Drive AI Boom

India is entering a new age of industrialization, as AI transforms how the world...

17/02/2026

India Fuels Its AI Mission With NVIDIA

India is the nexus of AI innovation this week as the host of the AI Impact Summit, which brings together global heads of state and industry to chart the future ...

16/02/2026

New Data Shows NVIDIA Blackwell Ultra Delivers up to 50x Better Performance and 35x Lower Costs for Agentic AI

The NVIDIA Blackwell platform has been widely adopted by leading inference provi...

12/02/2026

NVIDIA DGX Spark Powers Big Projects in Higher Education

At leading institutions across the globe, the NVIDIA DGX Spark desktop supercomputer is bringing data center class AI to lab benches, faculty offices and studen...

12/02/2026

Leading Inference Providers Cut AI Costs by up to 10x With Open Source Models on NVIDIA Blackwell

A diagnostic insight in healthcare. A character's dialogue in an interactive...

12/02/2026

GeForce NOW Turns Screens Into a Gaming Machine

The GeForce NOW sixth-anniversary festivities roll on this February, continuing a monthlong celebration of NVIDIA's cloud gaming service. This week brings ...

05/02/2026

GeForce NOW Celebrates Six Years of Streaming With 24 Games in February

Break out the cake and green sprinkles - GeForce NOW is turning six. Since launch, members have streamed over 1 billion hours, and the party's just getting...

04/02/2026

Nemotron Labs: How AI Agents Are Turning Documents Into Real-Time Business Intelligence

Editor's note: This post is part of the Nemotron Labs blog series, which exp...

03/02/2026

Everything Will Be Represented in a Virtual Twin, NVIDIA CEO Jensen Huang Says at 3DEXPERIENCE World

At 3DEXPERIENCE World in Houston, NVIDIA founder and CEO Jensen Huang and Dassau...

29/01/2026

Mercedes-Benz Unveils New S-Class Built on NVIDIA DRIVE AV, Which Enables an L4-Ready Architecture

Mercedes-Benz is marking 140 years of automotive innovation with a new S-Class b...

29/01/2026

Into the Omniverse: Physical AI Open Models and Frameworks Advance Robots and Autonomous Systems

Editor's note: This post is part of Into the Omniverse, a series focused on ...

29/01/2026

GeForce NOW Brings GeForce RTX Gaming to Linux PCs

Get ready to game - the native GeForce NOW app for Linux PCs is now available in beta, letting Linux desktops tap directly into GeForce RTX performance from the...

28/01/2026

Accelerating Science: A Blueprint for a Renewed National Quantum Initiative

Quantum technologies are rapidly emerging as foundational capabilities for economic competitiveness, national security and scientific leadership in the 21st cen...

22/01/2026

NVIDIA DRIVE AV Raises the Bar for Vehicle Safety as Mercedes-Benz CLA Earns Top Euro NCAP Award

AI-powered driver assistance technologies are becoming standard equipment, funda...

22/01/2026

Flight Controls Are Cleared for Takeoff on GeForce NOW

The wait is over, pilots. Flight control support - one of the most community-requested features for GeForce NOW - is live starting today, following its announce...

22/01/2026

From Pilot to Profit: Survey Reveals the Financial Services Industry Is Doubling Down on AI Investment and Open Source

AI has taken center stage in financial services, automating the research and exe...

22/01/2026

How to Get Started With Visual Generative AI on NVIDIA RTX PCs

AI-powered content generation is now embedded in everyday tools like Adobe and Canva, with a slew of agencies and studios incorporating the technology into thei...

21/01/2026

Largest Infrastructure Buildout In Human History': Jensen Huang on AI's Five-Layer Cake' at Davos

From skilled trades to startups, AI's rapid expansion is the beginning of th...

21/01/2026

Largest Infrastructure Buildout In Human History: Jensen Huang on AI's Five-Layer Cake at Davos

From skilled trades to startups, AI's rapid expansion is the beginning of th...

15/01/2026

Survive the Quarantine Zone and More With Devolver Digital Games on GeForce NOW

NVIDIA kicked off the year at CES, where the crowd buzzed about the latest gaming announcements - including the native GeForce NOW app for Linux and Amazon Fire...

13/01/2026

CEOs of NVIDIA and Lilly Share Blueprint for What Is Possible' in AI and Drug Discovery

NVIDIA and Lilly are putting together a blueprint for what is possible in the f...

09/01/2026

NVIDIA Unveils Multi-Agent Intelligent Warehouse and Catalog Enrichment AI Blueprints to Power the Retail Pipeline

Every that was easy shopping moment is made possible by teams working to hit s...

08/01/2026

Japan Science and Technology Agency Develops NVIDIA-Powered Moonshot Robot for Elderly Care

The next universal technology since the smartphone is on the horizon - and it ma...

08/01/2026

AI Copilot Keeps Berkeley's X-Ray Particle Accelerator on Track

In the rolling hills of Berkeley, California, an AI agent is supporting high-stakes physics experiments at the Advanced Light Source (ALS) particle accelerator....

08/01/2026

More Ways to Play, More Games to Love - GeForce NOW Wraps CES With Linux Support, Fire TV App, Flight Stick Controls

NVIDIA is wrapping up a big week at the CES trade show with a set of GeForce NOW...

07/01/2026

From Warehouse to Wallet: New State of AI in Retail and CPG Survey Uncovers How AI Is Rewiring Supply Chains and Customer Experiences

AI has transformed retail and consumer packaged goods (CPG) operations, enhancin...

05/01/2026

NVIDIA Expands Global DRIVE Hyperion Ecosystem to Accelerate the Road to Full Autonomy

At the CES trade show running this week in Las Vegas, NVIDIA announced that the ...

05/01/2026

NVIDIA DGX Spark and DGX Station Power the Latest Open-Source and Frontier Models From the Desktop

Open-source AI is accelerating innovation across industries, and NVIDIA DGX Spar...