Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

North America Stories

23/02/2026

Dignity Health Sports Park Improves Fan Engagement Through Daktronics LED Display, Control System Upgrade

Dignity Health Sports Park, a multi-use sports complex in Carson, Calif., and Da...

23/02/2026

Gotham Sports Cuts Prices for Streaming Packages

The Gotham Sports App, the exclusive direct-to-consumer streaming home of MSG Networks and the YES Network, is introducing more choice and value with new packag...

23/02/2026

Cobalt ARIA AUD-MON Receives Best of Show Awards at ISE 2026

Cobalt Digital, a designer and manufacturer of ST 2110 and SDI signal processing products, and a founding partner in the openGear initiative, announces that its...

23/02/2026

Telos Alliance Releases Nielsen Watermark Encoder Update for Linear Acoustic AERO TV Processors

Telos Alliance, which has specialized in broadcast audio for more than three dec...

23/02/2026

L3Harris Advances Powder-in, Engine-out Hypersonic Propulsion Manufacturing

L3Harris hypersonics concept illustration....

23/02/2026

L3Harris Celebrates Engineers Week 2026

By combining their innovative spirit and technical expertise, L3Harris engineers are solving our customers most urgent challenges and shaping a better tomorrow....

23/02/2026

Broadcasters Foundation of America Announces 2026 Honorees

Share Copy link Facebook X Linkedin Bluesky Email...

23/02/2026

DirecTV Adds Apple TV's Formula 1 Racing

Share Copy link Facebook X Linkedin Bluesky Email...

23/02/2026

Netflix Greenlights New Romcom Film Messily Ever After' (WT), A Decade-Long Relationship on a Wild Emotional Ride

Back to All News Netflix Greenlights New Romcom Film Messily Ever After' (...

23/02/2026

Netflix Announces 1 Million Donation To National Film And Television School

Back to All News Netflix Announces £1 Million Donation To National Film And Television School Social Impact 23 February 2026 GlobalUnited Kingdom Link copi...

23/02/2026

NVIDIA Brings AI-Powered Cybersecurity to World's Critical Infrastructure

As technologies and systems become more digitalized and connected across the world, operational technology (OT) environments and industrial control systems (ICS...

21/02/2026

OBS CTO Sotiris Salamouris on the Legacy of the 2026 Winter Games

With Software Defined Broadcasting more established in Milan Cortina look for Los Angeles 2028 to have less hardware and more cloud-based software systems...

21/02/2026

Invisible String: NBC Olympics' Marsha Bird on the Hidden Work That Makes the Games Shine

The SVP of Olympic Operations on turning CAD drawings into reality, building tru...

21/02/2026

NAB Urges FCC to Tamp Down Reallocation Plans for Upper C-Band

Share Copy link Facebook X Linkedin Bluesky Email...

21/02/2026

Chyron Paint 10.3 Adds New Visualization Features for Live Action

Share Copy link Facebook X Linkedin Bluesky Email...

21/02/2026

Netflix Unveils the Trailer of Accused', A Psychological Drama on Power, Perception and Public Judgment

Back to All News Netflix Unveils the Trailer of Accused', A Psychological ...

20/02/2026

Gravity Media, Green Couch Entertainment Partner to Create Original Programming for International Broadcasters

Gravity Media and Los Angeles-based Green Couch Entertainment announce a strateg...

20/02/2026

IMAX, Apple TV Bring 2026 FIA Formula One World Championship Races to Selected U.S. Locations

IMAX announces it is working with Apple TV to bring the 2026 FIA Formula One Wor...

20/02/2026

Big Upgrade to Spring Training Experience With New Daktronics Displays at Phillies' BayCare Ballpark

Daktronics has partnered with the Philadelphia Phillies to design, manufacture, ...

20/02/2026

ESPN To Launch Women's Sports Sundays, a New Summer Weekly Primetime Franchise

ESPN announces the upcoming launch of Women's Sports Sundays - a first-of-it...

20/02/2026

Sennheiser Spectera Handheld Makes Broadcast Debut at Super Bowl LX

As the Seattle Seahawks and New England Patriots faced off in the NFL's biggest sporting event of the season on Sun., Feb. 8, Sennheiser wireless solutions ...

20/02/2026

ESPN To Stream More MLB Spring Training Games Than Ever in 2026

ESPN announces its 2026 Major League Baseball spring training schedule, which includes four national games on ESPN, six games on ESPN Unlimited, and more than 2...

20/02/2026

Open Broadcast Systems Launches 200 Gigabit Ethernet

Open Broadcast Systems, which specializes in software-based professional video transport, has added support for 200 Gigabit Ethernet to its range of encoders an...

20/02/2026

Chyron PAINT 10.3 Adds New Ways To Visualize Live Sports

Chyron announces the release of PAINT 10.3, which is designed to help analysts and operators turn live action into clearer, faster on-air storytelling. PAINT 1...

20/02/2026

Live Spring Training Games Begin Feb. 20 on MLB Network With Yankees Against Orioles

With full squad workouts underway, MLB Network's live Spring Training game s...

20/02/2026

MLS Kickoff 2026: As FIFA Men's World Cup Looms, League's Media Operations and Apple TV Aim To Enhance Viewing Experience

Tech enhancements, marquee productions are expected to take advantage of a summe...

20/02/2026

SVG GameDay, Ep. 4: Temple Athletics' Paige Wisehaupt - Driving Digital Content for the Cherry & White

In-venue and creative video staffers at the professional and collegiate level ha...

20/02/2026

BBC Sport's Focus on Clips, Social Media and a Highlights-Free Games

Speaking with SVG Europe after one of Team GB's greatest days at a Winter Olympics, BBC Sport's head of major events, Ron Chakraborty, explains the broa...

20/02/2026

Warner Bros. Discovery on Making Olympic Magic for Multiple Local Markets From Its Studios in Cortina and Livigno

Making Winter Games Olympic magic is the goal for every broadcaster in Italy cov...

20/02/2026

OBS Curling Directors Brie Robertson and Susan Young Talk Olympic Experiences and Keeping the Story Going with the Lights Out

Curling, one of the least-dangerous Winter Olympic sports, is dominating the Mil...

20/02/2026

How Digital First' Goals Are Bringing Together Linear and Online Editorial for BBC Sport Coverage at the Winter Olympics

BBC Sport's presence at the 2026 Winter Games is centred around a significan...

20/02/2026

Pinning Digital Content Onto the Backbone of a Strong Linear Infrastructure for BBC Sport in Cortina

BBC Sport is bringing together its linear TV and streaming digital arms in a str...

20/02/2026

BBC Sport Talks Influencers and Athletes in its Bid to Pivot to a Digital First' Strategy for the Games

To broaden the appeal of winter sports at Milano Cortina, the BBC has integrated...

20/02/2026

DIRECTV Adds Apple TV's Formula 1 Racing for Residential and DIRECTV FOR BUSINESS Customers

Just in time for the start of Apple TV's inaugural season as the exclusive U...

20/02/2026

NBC Olympics VP Creative David Barton Dives into Milano-Cortina 2026 Graphics Package, Now You Know' Animations

One big challenge was to depict the character of each of very different and wide...

20/02/2026

By Design Director Asks Festival Audience to Flirt With Their Chairs at Premiere

(L-R) Writer-director Amanda Kramer photographs the photographers at the premiere of her film By Design at the Library Center Theatre in Park City. (Photo by ...

20/02/2026

Super Bowl LX Delivers 125.6 Million Viewers

NEW YORK - February 10, 2026 - An estimated 125.6* million viewers watched Super Bowl LX on Sunday, February 8, according to Nielsen's Big Data Panel meas...

20/02/2026

Nielsen: Super Bowl LX Final Viewership Rises to 125.6 Million Viewers with Verified Big Data

NEW YORK - February 19, 2026 - Nielsen today shared updated and final Super Bowl...

20/02/2026

NHPBS Taps Heartland Video Systems for ATSC 3.0 Launch

Share Copy link Facebook X Linkedin Bluesky Email...

20/02/2026

Lightware and Cisco interoperability brings AV system suc...

A leading global investment bank, with offices at Two International Finance Centre in Hong Kong, partnered with systems integrators Global Vision Engineering (G...

20/02/2026

Rise AV and Rise Broadcast unite for International Womens...

Rise AV and Rise Broadcast, the global not-for-profit organisations dedicated to improving gender diversity across technical industries, have today announced a ...

20/02/2026

Open Broadcast Systems launches Two Hundred Gigabit Ether...

Open Broadcast Systems, the leader in software-based professional video transport, has added support for 200 Gigabit Ethernet to its range of encoders and decod...

20/02/2026

Signiant Launches Customer Advisory Board to Help Shape t...

Signiant today announced the formation of its Customer Advisory Board (CAB), bringing together a select group of customers to collaborate on product strategy, r...

20/02/2026

PTZOptics launches its Visual Reasoning initiative and pa...

PTZOptics today announced the launch of its Visual Reasoning initiative that makes video more actionable by combining robotic PTZ camera systems, AI, and open i...

20/02/2026

DELTA and Amino Complete Certification of Amigo 7N Androi...

Amino, a global media technology provider delivering devices, software and cloud services that simplify and elevate video delivery, today announced the successf...

20/02/2026

SMPTE Opens Call for Papers for 2026 Media Technology Sum...

SMPTE , the home of media professionals, technologists, and engineers, today announced its call for technical papers for the SMPTE 2026 Media Technology Summit....

20/02/2026

Granicus Standardizes Hybrid Government-Grade Video Infra...

Wowza Media Systems today announced that Granicus, a leading provider of digital engagement solutions for governments, continues to rely on Wowza to power its h...

20/02/2026

CBS Baltimore Launches New AR/VR Studio

Share Copy link Facebook X Linkedin Bluesky Email...

20/02/2026

IAB Tech Lab Opens Public Comment on Live Event Ad Playbook

Share Copy link Facebook X Linkedin Bluesky Email...