Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

28/04/2026

Sound for the NFL Draft Show Balances Often Conflicting Goals

The audio team for the entertainment event must blend speech intelligibility with full-range music reproduction while considering the broadcast Last week's...

28/04/2026

Pac-12 Releases New Primary Mark Ahead of 2026-27 Season

The Pac-12 Conference has released an updated primary mark and logo as the starting point of the new league's brand identity. The mark was soft-launched acr...

28/04/2026

DP World Tour to Become First Professional Sports Organization to Use Amazon Leo Satellite Connectivity

The DP World Tour and Amazon Leo have signed an agreement making Amazon's lo...

28/04/2026

Pixellot and HELIOS Announce Automated Hockey Shift Video Integration

Pixellot and HELIOS have announced an integration that automatically converts full-game hockey video into individualized shift videos for each athlete, without ...

28/04/2026

Daktronics Installs New LED Video Display at Asheville Tourists Ballpark

Daktronics has partnered with the Asheville Tourists to manufacture and install a new LED video display. The installation was completed in late 2025 and is now ...

28/04/2026

Eutelsat and PCTV Renew Long-Term Video Distribution Partnership in Mexico

Eutelsat has announced the renewal of its partnership with PCTV, a content aggregation and distribution company in Mexico and part of Megacable Holdings, for co...

28/04/2026

Gary SouthShore RailCats Install New Daktronics LED Display at U.S. Steel Yard

Daktronics has partnered with the Gary SouthShore RailCats to install a new LED video display at U.S. Steel Yard, replacing the previous Daktronics display inst...

28/04/2026

Telos Alliance and College Radio Foundation Award Omnia.11 Processor to Wright State Universitys WWSU-FM

Telos Alliance and the College Radio Foundation have announced that WWSU-FM of W...

28/04/2026

Inside the Mix: A1s Florian Brown and James Deason on Golfs Biggest Broadcasts

Golf viewership is growing. The 2025 Ryder Cup drew five million viewers in the UK, a 45% increase over the 2023 event. The US Open was the most streamed golf e...

28/04/2026

The CW Network Acquires Exclusive Broadcast Rights to WWE NXT Premium Live Events

The CW Network and WWE, part of TKO Group Holdings (NYSE: TKO), have announced t...

28/04/2026

NAB 2026: AIMS Wins NAB Show 2026 Product of the Year Award for IPMX

The Alliance for IP Media Solutions (AIMS) has announced that the Internet Protocol Media Experience (IPMX) suite of standards and specifications has been named...

28/04/2026

NAB 2026 SportsTechBuzz in Review: A Look Back at the Big News From 150+ Companies in Vegas

The 2026 NAB Show is in the books and the show once again served up a cavalcade ...

28/04/2026

Gray Media, RAJ Sports Launch New Womens Sports RSN in Portland - Rose City SportsNet

Gray Media and RAJ Sports have announced Rose City SportsNet (RCSN), a new netwo...

28/04/2026

Spotify Reports First Quarter 2026 Earnings

Today, we announced our First Quarter 2026 earnings, starting the Year of Raising Ambition with strong momentum across the business and continued innovation acr...

28/04/2026

Spotify rapporterar resultat fr frsta kvartalet 2026

I dag presenterade vi v rt resultat f r det f rsta kvartalet 2026. Vi inleder ret med starkt momentum i hela verksamheten och fortsatt innovation p plattforme...

28/04/2026

Mojave introduce the MA-C vocal mic

New handheld promises studio performance for the stage Mojave have just introduced a new live-focused handheld vocal mic created by award-winning designer D...

28/04/2026

VSTOPIA announce Dynamic Split Module for Ableton

Max for Live device offers AI-powered stem separation Dynamic Split Module (DSM) is a new Max for Live device created by Ostin Solo, a developer and musican...

28/04/2026

Focusrite unveil the ISA C8X

First interface equipped with ISA preamps Focusrite have just announced the launch of a new high-end audio interface that features a pair of their legendary...

28/04/2026

Nielsen and Triton Digital Collaborate to Bring Greater Visibility to Podcast Audiences in Nielsen's Media Impact Tool

Triton Digital's Podcast Metrics Demos+ Data Integration Enables Comprehensi...

28/04/2026

Paramount Skydance Will be 49.5% Foreign Owned After WBD Merger

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

KEET PBS Deploys PMVG TechBundle Services To Modernize Operations

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

TAG Video Systems Lens Wins Industry Awards for Cutting T...

TAG Video Systems, the leading IP-native Realtime Media Platform, today announced that Lens, its visual service health interface for broadcast operations, recei...

28/04/2026

AIMS Wins NAB Show 2026 Product of the Year Award for IPM...

Open AV-over-IP Standard Recognized in IT Networking/Infrastructure and Security Category The Alliance for IP Media Solutions (AIMS) today announced that the ...

28/04/2026

VFX History: Slit Scan

VFX History: Slit Scan Graham Quince April 28, 2026 0 Comments How did 2001: A Space Odyssey, Star Wars, Doctor Who and Star Trek: The Next Generation...

28/04/2026

These DaVinci Resolve Effects Will Make You a More Creative Colorist

These DaVinci Resolve Effects Will Make You a More Creative Colorist Kasia Jarco April 28, 2026 0 Comments Creativity in color grading is not about ha...

28/04/2026

A Simple Introduction to Cavalry: Indexed Circle

A Simple Introduction to Cavalry: Indexed Circle Simon Ubsdell April 28, 2026 0 Comments In this new introductory tutorial for Cavalry we're going...

28/04/2026

Rise Upskill - Applications Now Open for Free Global Trai...

Rise, the award-winning advocacy group for gender diversity in the broadcast and media technology sector, is pleased to announce a new global training programme...

28/04/2026

Clear Com Announces New Roles for Brian Grahn and Ben Tur...

Clear-Com has appointed Brian Grahn as Market Outreach Manager of the Americas and Ben Turnwell as Business Development Manager for EMEA live, expanding their ...

28/04/2026

LiveU Steps into the Future at MPTS 2026 with the Introdu...

LiveU is inviting MPTS visitors to step into the companys new Q Era on Stand D32, at The Grand Hall, Olympia, London (May 13-14). The company will showcase its ...

28/04/2026

IBC launches 2026 Innovation Awards to spotlight real-wor...

IBC today announces the launch of the IBC2026 Innovation Awards, with nominations now open for projects, programmes and initiatives that exemplify breakthrough ...

28/04/2026

WNBA to Stream All Preseason Games for Free

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

Nexstar Media Charitable Foundation Sets 30 Days of Giving'

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

Sinclair's Chief Compliance Officer Jeff Lewis to Retire

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

Nielsen Introduces Predictive Sales Lift' Tool

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

Sencore's VB440 Monitoring, Analysis Tool Debuts at NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

Pinterest Makes a Major Push into CTV Advertising

Share Copy link Facebook X Linkedin Bluesky Email...

28/04/2026

Introducing Nx 3-Strip v2 - A Physics-Based Technicolor Reconstruction for DaVinci Resolve

Introducing Nx 3-Strip v2 - A Physics-Based Technicolor Reconstruction for DaVin...

28/04/2026

Tribeca Festival Marks 25 Years With Star-Studded Talks, Reunions & Retrospectives

April 28th, 2026 Press Materials Available Here TRIBECA FESTIVAL MARKS 25 YEAR...

28/04/2026

Sky Documentaries announces Savage Mountain, the controversial true story behind Kristin Harilas record-breaking 14 Peaks climb

In 2023, Norwegian climber Kristin Harila set out to break a mountaineering reco...

28/04/2026

LinkedIn Top Companies 2026: Where Career Growth Is...

LinkedIn Top Companies 2026: Where Career Growth Is Happening Now Published on Apr 28, 2026 Categories: Data and insights LinkedIn Corporate Communication...

28/04/2026

Into the Omniverse: Manufacturing's Simulation-First Era Has Arrived

Editor's note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners, and enterprises can transform their workflows ...

28/04/2026

NVIDIA Launches Nemotron 3 Nano Omni Model, Unifying Vision, Audio and Language for up to 9x More Efficient AI Agents

AI agent systems today juggle separate models for vision, speech and language - ...

28/04/2026

RT NEWS ANNOUNCES SEAN WHELAN AS NEW LONDON CORRESPONDENT

RT News is pleased to announce the appointment of Sean Whelan as its new London Correspondent. Sean has held the role of Washington Correspondent for the last...

28/04/2026

Six well-known figures delve into the 1926 Census archive

Joseph O'Connor, Eileen Walsh, Louise Duffy, Mick Lynch, Gormfhlaith N Thuairisg and Dermot Bannon take a personal look at life 100yrs ago in new TV docume...