Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

North America Stories

11/04/2026

Federal Judge Extends Nexstar/Tegna TRO, Softens Some Provisions

Share Copy link Facebook X Linkedin Bluesky Email...

11/04/2026

Sling TV Launches $19.99 a Month Sling Essentials with ESPN

Share Copy link Facebook X Linkedin Bluesky Email...

11/04/2026

NAB Show Launches Content Creator VIP Program

Share Copy link Facebook X Linkedin Bluesky Email...

11/04/2026

FCC Announces Tentative Agenda for April Open Meeting

Share Copy link Facebook X Linkedin Bluesky Email...

11/04/2026

GARR and Cubbit launch the first geo-distributed storage...

Pilot phase begins for a new national infrastructure designed to safeguard academic and research data with full local data control, sovereignty, resilience, and...

11/04/2026

How to Stream Coachella 2026 at Home

How to Stream Coachella 2026 at Home Check this years stacked schedule for the annual music festivals full lineup, including when Berklee artists from Laufey ...

11/04/2026

Time Travel

Time Travel As Berklee on the Road programs in Puerto Rico and Italy mark decades-long anniversaries, we journey into the past and step into the future. Apri...

11/04/2026

April 10, 2026

Improving vaccine design for Ebola, HIV and more Scripps Research scientists and colleagues develop a nanodisc platform that offers a clearer view of how key vi...

10/04/2026

The Invisible OPEX Killer: Is Your Server Room Dragging You Down?

The Invisible OPEX Killer: Is Your Server Room Dragging You Down? In the broadcast world, we talk a lot about uptime. We talk about talent retention, latency...

10/04/2026

NAB 2026: Imagine Communications to Showcase Expanded Multiviewer Portfolio

Imagine Communications will showcase its multiviewer portfolio at NAB Show 2026 (April 19-22, Booth N1328, Las Vegas Convention Center), including Prismon and t...

10/04/2026

NAB 2026: Chyron Releases PRIME VSAR 2.3 with Updated Unreal Engine Integration

Chyron has released PRIME VSAR 2.3, an update to its virtual set and augmented reality solution for broadcast. The release adds compatibility with Unreal Engine...

10/04/2026

NAB 2026: Techex to Showcase New tx darwin Capabilities

Techex will exhibit at NAB Show 2026 (Booth W2267, April 19-23, Las Vegas Convention Center), demonstrating new tx darwin features including consumer multiview,...

10/04/2026

NAB 2026: NDI to Showcase Ecosystem and NDI 6.3

NDI will exhibit at NAB Show 2026, demonstrating its IP video ecosystem through live partner integrations, NDI 6.3 features, AI metadata workflows, and creator ...

10/04/2026

FOR-A Acquires Tamura Corporations Information Equipment Business

FOR-A has announced the acquisition of all shares of Tamu Radiance Corporation, a new company spun off from the Information Equipment Business of Tamura Corpora...

10/04/2026

NAB 2026: InSync Technology to Unveil New Video Processing and Frame Rate Conversion Products

InSync Technology will showcase new and updated video conversion products at NAB...

10/04/2026

TNT Sports and DAZN Announce Monthly Boxing Event Series in the United States

TNT Sports and DAZN have announced a partnership to air monthly boxing events in the United States under the brand The Fight. The series will be promoted in p...

10/04/2026

Panasonic Introduces SQ3 Series 4K LCD Displays for Professional Environments

Panasonic Projector and Display has announced the SQ3 Series of 4K LCD displays as part of its MEVIX professional display portfolio. All sizes will be available...

10/04/2026

Amagi Adds Agentic Capabilities to Its Media Operations Platform

Amagi has announced the addition of Agentic Media Operations to its Amagi NOW platform, integrating AI reasoning agents across its media supply chain workflows ...

10/04/2026

LTN Announces Network Enhancements Ahead of C-Band Spectrum Auction

LTN has announced enhancements to its global IP video network targeting broadcasters transitioning from satellite distribution. The updates come ahead of US fed...

10/04/2026

Daktronics Installs New LED Displays at Yankee Stadium

Daktronics has installed new LED displays at Yankee Stadium, upgrading the main centerfield board, two flanking boards, and two ribbon displays spanning the 200...

10/04/2026

NAB 2026: Harmonic Announces AI and Cloud Updates to Hybrid Streaming Solution

Harmonic has announced updates to its hybrid streaming solution, including Model Context Protocol (MCP) connectivity for AI applications, cloud-native deploymen...

10/04/2026

NAB 2026: MultiDyne to Debut FiberSaver-10G and VF-9100

MultiDyne Video and Fiber Optic Systems will introduce two new fiber transport products at NAB Show 2026 (Booth C4425, April 19-22): the FiberSaver-10G waveleng...

10/04/2026

NAB 2026: Telos Alliance and ip-studio to Demonstrate STUDIO ZERO

Telos Alliance and ip-studio will demonstrate STUDIO ZERO, a cloud-hosted virtual studio, at NAB Show 2026. First introduced at NAB Show 2023, STUDIO ZERO integ...

10/04/2026

Pixotope and d&b Solutions Announce Strategic Partnership for XR and Virtual Studio Production

d&b solutions, a London-based audio-visual, lighting, and media integration grou...

10/04/2026

ARRI and SmallHD Announce Lens Data Monitor Overlay License for Hi-5 and Hi-5 SX

ARRI and SmallHD have announced a new expansion license for ARRI's Hi-5 and Hi-5 SX hand units that displays lens data overlays on supported SmallHD monitor...

10/04/2026

Roku to Stream Exclusive Savannah Bananas Game Package on Roku Sports Channel

Roku and the Banana Ball Championship League (BBCL) have announced an exclusive streaming partnership to bring five BBCL games to the Roku Sports Channel in 202...

10/04/2026

Ratings Roundup: More Than 18 Million Fans Tune Into 2026 NCAA Mens March Madness on TNT and CBS Sports

Ratings Roundup is a rundown of recent rating news and is derived from press rel...

10/04/2026

No Other Land, Mr. Nobody Against Putin,and More Sundance Institute-Supported Films Nominated for Peabody Awards

The Peabody Awards don't just recognize great storytelling, they spotlight t...

10/04/2026

2026 NAB Show Exhibitor Insight: Bitcentral

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Bitcentral To Feature Connected Media Workflows At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Bitcentral to Showcase Connected Media Workflows and Inte...

Bitcentral, a leading provider of professional media solutions for broadcast and digital video, will showcase its latest innovations at NAB Show 2026 (Booth W28...

10/04/2026

Ikegami to Introduce Expanded Range of Broadcast Production Solutions at NAB 2026

Ikegami to Introduce Expanded Range of Broadcast Production Solutions at NAB 202...

10/04/2026

AJA Debuts SMPTE ST 2110 and openGear Solutions Ahead of NAB 2026

AJA Debuts SMPTE ST 2110 and openGear Solutions Ahead of NAB 2026 Brie Clayton April 10, 2026 0 Comments New gear and updates address evolving hybrid ...

10/04/2026

Portland Fire+ Streaming Platform Launches

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Tod Musgrave Joins Proton as U.S. Sales & Marketing Director

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Proton Expands Minicam Portfolio With Proton Pro At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

FCC To Vote on Changes to Audible Crawl Rule

Share Copy link Facebook X Linkedin Bluesky Email...

10/04/2026

Frequency Launches AI Platform for Streaming Television a...

Frequency, the engine behind the worlds leading streaming television channels, today launched its AI platform for Frequency Studio, powering the entire channel ...

10/04/2026

Jnger Audio Joins EBU ADM Integration Group as Founding Member to Help Advance ADM/S-ADM Integration

J nger Audio Joins EBU ADM Integration Group as Founding Member to Help Advance...

09/04/2026

NAB 2026: Zixi to Demonstrate Live Video Workflows and Satellite Replacement

Zixi will demonstrate IP-based live video workflow solutions at NAB Show 2026 (Booth W2057). The industry is moving quickly toward IP-based distribution as br...

09/04/2026

Deloitte Research: Women's Elite Sports Revenues Expected to Reach at Least $3 Billion in 2026

Global women's elite sports revenues are expected to reach at least $3 billi...

09/04/2026

Monitor Engineer Gavin Tempany Mixes Kylie Minogue's Tension Tour on Solid State Logic L550 Plus

Monitor engineer Gavin Tempany mixed Kylie Minogue s Tension Tour on a Solid Sta...

09/04/2026

NAB 2026: KOKUSAI DENKI Electric America to Debut New 4K Camera and Remote Control Panel

KOKUSAI DENKI Electric America will exhibit at NAB Show 2026 (Booth C5507), debu...

09/04/2026

NBC Sports Reviews Innovations and Milestones from Its 2025-26 NBA Regular Season

With the 2025-26 NBA regular season concluded and the playoffs beginning next we...

09/04/2026

NAB 2026: Telestream and Mimir Announce Integration for Ingest-to-Editorial Workflows

Telestream and Mimir have announced an integration connecting Telestream's V...

09/04/2026

NAB 2026: Bitmovin Expands Live Encoding and Observability Solutions for End-to-End Live Streaming Monitoring

Bitmovin has expanded its Live Encoding and Observability solutions to provide r...

09/04/2026

Nashville Predators and Scripps Sports Announce Multi-Year Broadcast Agreement

The Nashville Predators and Scripps Sports have announced a multi-year media rights agreement covering local preseason, regular season, and first-round playoff ...

09/04/2026

ASG Partners with Beam Dynamics for Asset Intelligence Platform

Advanced Systems Group, LLC has announced a partnership with Beam Dynamics to offer the Beam Asset and License Intelligence Platform to its clients. The platfor...

09/04/2026

NAB 2026: Lawo Introduces Edge One Converged Video and Audio Stagebox

Lawo has unveiled Edge One, a combined video and audio stagebox for broadcast and Pro AV workflows. The device will be on display at NAB Show (Booth C2108, Apri...

09/04/2026

NAB 2026: SMPTE to Host ST 2110 IP Media Roadshow

The Society of Motion Picture and Television Engineers (SMPTE) will host the SMPTE ST 2110 IP Media Roadshow on Tuesday, April 21, 2026, at the Las Vegas Conven...