Sony Pixel Power calrec Sony

TOPS of the Class: Decoding AI Performance on RTX AI PCs and Workstations

12/06/2024

Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, software, tools and accelerations for RTX PC users.

The era of the AI PC is here, and it's powered by NVIDIA RTX and GeForce RTX technologies. With it comes a new way to evaluate performance for AI-accelerated tasks, and a new language that can be daunting to decipher when choosing between the desktops and laptops available.

While PC gamers understand frames per second (FPS) and similar stats, measuring AI performance requires new metrics.

Coming Out on TOPS The first baseline is TOPS, or trillions of operations per second. Trillions is the important word here - the processing numbers behind generative AI tasks are absolutely massive. Think of TOPS as a raw performance metric, similar to an engine's horsepower rating. More is better.

Compare, for example, the recently announced Copilot+ PC lineup by Microsoft, which includes neural processing units (NPUs) able to perform upwards of 40 TOPS. Performing 40 TOPS is sufficient for some light AI-assisted tasks, like asking a local chatbot where yesterday's notes are.

But many generative AI tasks are more demanding. NVIDIA RTX and GeForce RTX GPUs deliver unprecedented performance across all generative tasks - the GeForce RTX 4090 GPU offers more than 1,300 TOPS. This is the kind of horsepower needed to handle AI-assisted digital content creation, AI super resolution in PC gaming, generating images from text or video, querying local large language models (LLMs) and more.

Insert Tokens to Play TOPS is only the beginning of the story. LLM performance is measured in the number of tokens generated by the model.

Tokens are the output of the LLM. A token can be a word in a sentence, or even a smaller fragment like punctuation or whitespace. Performance for AI-accelerated tasks can be measured in tokens per second.

Another important factor is batch size, or the number of inputs processed simultaneously in a single inference pass. As an LLM will sit at the core of many modern AI systems, the ability to handle multiple inputs (e.g. from a single application or across multiple applications) will be a key differentiator. While larger batch sizes improve performance for concurrent inputs, they also require more memory, especially when combined with larger models.

The more you batch, the more (time) you save. RTX GPUs are exceptionally well-suited for LLMs due to their large amounts of dedicated video random access memory (VRAM), Tensor Cores and TensorRT-LLM software.

GeForce RTX GPUs offer up to 24GB of high-speed VRAM, and NVIDIA RTX GPUs up to 48GB, which can handle larger models and enable higher batch sizes. RTX GPUs also take advantage of Tensor Cores - dedicated AI accelerators that dramatically speed up the computationally intensive operations required for deep learning and generative AI models. That maximum performance is easily accessed when an application uses the NVIDIA TensorRT software development kit (SDK), which unlocks the highest-performance generative AI on the more than 100 million Windows PCs and workstations powered by RTX GPUs.

The combination of memory, dedicated AI accelerators and optimized software gives RTX GPUs massive throughput gains, especially as batch sizes increase.

Text-to-Image, Faster Than Ever Measuring image generation speed is another way to evaluate performance. One of the most straightforward ways uses Stable Diffusion, a popular image-based AI model that allows users to easily convert text descriptions into complex visual representations.

With Stable Diffusion, users can quickly create and refine images from text prompts to achieve their desired output. When using an RTX GPU, these results can be generated faster than processing the AI model on a CPU or NPU.

That performance is even higher when using the TensorRT extension for the popular Automatic1111 interface. RTX users can generate images from prompts up to 2x faster with the SDXL Base checkpoint - significantly streamlining Stable Diffusion workflows.

ComfyUI, another popular Stable Diffusion user interface, added TensorRT acceleration last week. RTX users can now generate images from prompts up to 60% faster, and can even convert these images to videos using Stable Video Diffuson up to 70% faster with TensorRT.

TensorRT acceleration can be put to the test in the new UL Procyon AI Image Generation benchmark, which delivers speedups of 50% on a GeForce RTX 4080 SUPER GPU compared with the fastest non-TensorRT implementation.

TensorRT acceleration will soon be released for Stable Diffusion 3 - Stability AI's new, highly anticipated text-to-image model - boosting performance by 50%. Plus, the new TensorRT-Model Optimizer enables accelerating performance even further. This results in a 70% speedup compared with the non-TensorRT implementation, along with a 50% reduction in memory consumption.

Of course, seeing is believing - the true test is in the real-world use case of iterating on an original prompt. Users can refine image generation by tweaking prompts significantly faster on RTX GPUs, taking seconds per iteration compared with minutes on a Macbook Pro M3 Max. Plus, users get both speed and security with everything remaining private when running locally on an RTX-powered PC or workstation.

The Results Are in and Open Sourced But don't just take our word for it. The team of AI researchers and engineers behind the open-source Jan.ai recently integrated TensorRT-LLM into its local chatbot app, then tested these optimizations for themselves.

Source: Jan.ai The researchers tested its implementation of TensorRT-LLM against the open-source llama.cpp inference engine across a variety of GPUs and CPUs used by the community. They found that TensorRT is 30-70% faster than llam
LINK: https://blogs.nvidia.com/blog/ai-decoded-tops/...
See more stories from nvidia

North America Stories

16/06/2026

Neumann MT 48 Receives Major Firmware 2.0 Update

Neumann.Berlin has released firmware version 2.0 for the MT 48 audio interface, adding plugin compatibility, expanded Dante networking options, broadcast encode...

16/06/2026

TVNewsCheck Opens Nominations for 2027 Women in Technology Awards

TVNewsCheck has announced that nominations are now open for its 2027 Women in Technology Awards, to be presented at NAB Show 2027 on Tuesday, April 6 in the Med...

16/06/2026

Clear-Com Introduces Avalon IP Intercom Platform

Clear-Com has announced Avalon, a 1RU IP intercom platform for broadcast, live events, and production environments. Designed for IP-only workflows, Avalon suppo...

16/06/2026

SNS EVO Enables Remote and Distributed Video Editing Workflows

SNS has published a guide to remote video editing workflows using its EVO shared storage platform and companion tools, covering use cases ranging from home edit...

16/06/2026

Richmond Flying Squirrels Deploy Grass Valley LDX 110 Cameras at CarMax Park

Grass Valley has announced that the Richmond Flying Squirrels, a Minor League Baseball affiliate of the San Francisco Giants, have deployed five Grass Valley LD...

16/06/2026

AIMS Launches Free Official IPMX Training Series Online

The Alliance for IP Media Solutions (AIMS) has announced the launch of the Official IPMX Training Series, a free online program covering the design, configurati...

16/06/2026

Swerve Womens Sports Announces Distribution Deals with Fubo, Plex, Amazon Fire TV, and Anoki AI

Swerve TV has announced distribution agreements with Fubo, Plex, Amazon Fire TV,...

16/06/2026

ATP and TikTok Expand Global Content Partnership

ATP and TikTok have announced an expansion of their global content partnership, extending the ATP's TikTok hub powered by TikTok GamePlan to cover all nine ...

16/06/2026

FOX Sports Turns Los Angeles Pico Lot Into Its FIFA World Cup Production Nerve Center

Network's LA facility serves as the heart of a sprawling operation built to ...

16/06/2026

300+ Records a Day, 150 TB Daily, and a Relentless Content Avalanche: Inside FOX Sports' World Cup Media Engine

At Pico, the network's media-management team is supporting a flood of HBS fe...

16/06/2026

NHL Games Leaving CBC in Canada as Sublicense With Rogers Sportsnet Ends

The NHL will no longer air on CBC after the pulic broadcasters and national rights-holder Rogers Sportsnet were unable to come to agreement. After a successfu...

16/06/2026

SVG New Sponsor Spotlight: Virtual Eye's Ben Taylor on Making Live Sports More Valuable and Entertaining Through Data-Driven Graphics

As live sports broadcasters continue to seek new ways to make complex action mor...

16/06/2026

Thats BRISK, Baby! FOX Sports' Broadcast Remote IP Studio Kits Bring World Cup Fan Energy Back to Pico

Built with the 2026 FIFA World Cup in mind, these small but mighty IP-based tran...

16/06/2026

Chyron Unveils Chyron Weather 2.4

Share Copy link Facebook X Linkedin Bluesky Email...

16/06/2026

Historic Zhuque-3 Reusable Rocket Test Mission Captured with URSA Cine Immersive

Historic Zhuque-3 Reusable Rocket Test Mission Captured with URSA Cine Immersive Brie Clayton June 16, 2026 0 Comments Apple Immersive Video puts view...

16/06/2026

SMPTE Plans ST 2110 Education Summer Programs

Share Copy link Facebook X Linkedin Bluesky Email...

16/06/2026

Rise Awards Returns for 2026 to Celebrate Excellence in B...

Rise WIB, the award-winning advocacy group championing gender diversity and career progression across the broadcast and media technology industry, today announc...

16/06/2026

Limecraft Expands its Media Production Platform with Team...

Limecraft today announced the availability of Limecraft 2026.4, the fourth of eight planned platform releases this year. The update introduces Team-Based Access...

16/06/2026

Perry Sook: Big Tech Poses 'Very Urgent Threat to Broadcast Stations

Share Copy link Facebook X Linkedin Bluesky Email...

16/06/2026

FIFA World Cup Delivers Record Ratings on Fox

Share Copy link Facebook X Linkedin Bluesky Email...

16/06/2026

AIMS Launches the Official IPMX Training Series Online

Free Program Supports IPMX Education from Foundational Concepts Through System and Network Design The Alliance for IP Media Solutions (AIMS) today announced t...

16/06/2026

Fastest, Largest, Strongest: NVIDIA Blackwell Sweeps MLPerf Training 6.0

Every breakthrough AI model starts the same way: with a training run. The infrastructure running those training jobs shapes everything: how fast teams can itera...

15/06/2026

University of South Carolina's Valerie Gerfin on Gamecock Productions' Growth, Upgrades at Williams-Brice Stadium

One of the more exciting internal video production divisions within a college at...

15/06/2026

Fox Corp. To Acquire Roku, Pairs Live Sports Powerhouse With Major CTV Platform

The deal valued at $22 Billion is expected to close in the first half of 2027...

15/06/2026

Golf Channel Mobile to Live Stream 2026 Arnold Palmer Cup Beginning July 13th

Golf Channel and the Arnold Palmer Cup have announced a partnership to livestream the 2026 Arnold Palmer Cup on Golf Channel Mobile and GolfChannel.com. The tou...

15/06/2026

TikTok and Panini Launch Digital Collectible Card Experience for FIFA World Cup 2026

TikTok and Panini have announced a partnership to bring a digital collectible ca...

15/06/2026

Cosm and Monster Energy Launch First Full-Dome Immersive Advertisement in Shared Reality Venues

Cosm and Monster Energy have announced the debut of the first full-dome immersiv...

15/06/2026

Fox Nation and Real American Freestyle Sign International Media Rights Deal

Real American Freestyle (RAF) and Fox Nation have announced an exclusive streaming agreement for three RAF international events, beginning with RAF Georgia on J...

15/06/2026

FanConnect and Extreme Networks Announce IPTV Integration for Large Venue Deployments

FanConnect has announced a partnership with Extreme Networks integrating FanConn...

15/06/2026

2026 Sundance Institute Ignite x Adobe Fellows Named

Ten Emerging Filmmakers Ages 18 to 25 Will Start Fellowship Year at Ignite Lab from June 14-19 LOS ANGELES, CA, June 15, 2026 - The nonprofit Sundance Institut...

15/06/2026

Clear-Com Introduces Avalon IP Intercom Platform

Share Copy link Facebook X Linkedin Bluesky Email...

15/06/2026

DoJ Approves Paramount Skydance, Warner Bros. Discovery Merger

Share Copy link Facebook X Linkedin Bluesky Email...

15/06/2026

Clear-Com Introduces Avalon IP Station for Modern Communi...

Clear-Com has introduced Avalon , a purpose built 1RU IP intercom communication platform for modern networked production, designed to simplify and scale workfl...

15/06/2026

Fox Makes CTV Play with Roku Acquisition

Share Copy link Facebook X Linkedin Bluesky Email...

15/06/2026

Gray Announces Plans to Expand Lansing, Mich. Broadcast HQ

Share Copy link Facebook X Linkedin Bluesky Email...

15/06/2026

Richmond Flying Squirrels Raise the Bar for Live Baseball...

MiLB Club Deploys LDX 110 Cameras at CarMax Park to Deliver A New Standard in Engaging Fan Experience Grass Valley today announced that the Richmond Flying Sq...

15/06/2026

Detach from Direct-Attached: How Remote Editing with EVO Keeps Creative Teams Moving

Detach from Direct-Attached: How Remote Editing with EVO Keeps Creative Teams Mo...

14/06/2026

HBO Comedy Rooster Shot with URSA Cine 17K 65

HBO Comedy Rooster Shot with URSA Cine 17K 65 Brie Clayton June 14, 2026 0 Comments Large format brings viewers intimately close to characters. Black...

13/06/2026

Spectrum Reach Taps Anoki AI for Contextual Intelligence

Share Copy link Facebook X Linkedin Bluesky Email...

13/06/2026

Google TV Launches Soccer Hub, New Voice Command Features

Share Copy link Facebook X Linkedin Bluesky Email...

12/06/2026

YES Network and Gotham Sports App to Air Seven Athletes Unlimited Softball League Games

YES Network and The Gotham Sports App will air seven Athletes Unlimited Softball...

12/06/2026

UFL to Feature FAST Innovation Suite at 2026 United Bowl

The United Football League will host its FAST Innovation Suite at the 2026 United Bowl presented by Credit One Bank on Saturday, June 13 at 3:00 p.m. ET at Audi...

12/06/2026

InfoComm 2026: PTZOptics and LayerJot to Demo AI-Driven Camera Control

PTZOptics and LayerJot will present live demonstrations at InfoComm 2026 showing how natural-language AI prompting, robotic camera control, and on-device comput...

12/06/2026

InfoComm 2026: MultiDyne to Debut VF-9100 Fiber Transport Platform and Crescendo Audio Monitor

MultiDyne Video and Fiber Optic Systems will exhibit at InfoComm 2026, featuring...

12/06/2026

Eurovision Services Deploys Ateme Software-Based Frame-Rate Conversion

Ateme has announced that Eurovision Services is using Ateme's software-based frame-rate conversion technology for international live event workflows. The de...

12/06/2026

Bitmovin, Simplestream, and Xperi Partner to Support OTT Services on TiVo OS

Bitmovin and Simplestream have announced a partnership with Xperi to simplify the launch of OTT streaming services on TiVo OS smart TVs and devices. The collabo...

12/06/2026

Net Insight Deploys Nimbra 520 and Nimbra Edge for Multinational Corporate Live Production Workflow

Net Insight has announced that a multinational technology company is deploying a...

12/06/2026

MLB Players Inc., Athletes First Announce Content Partnership

MLB Players Inc., the business arm of the MLB Players Association, has announced a partnership with Athletes First to develop and sell brand partnerships across...

12/06/2026

G&D and VuWall Announce CommandKeyboard-Advanced for Network-Independent Control Room Operations

Guntermann and Drunck (G&D) and VuWall have announced the CommandKeyboard-Advanc...