Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

North America Stories

04/03/2026

Lega Basket Serie A Modernizes Media Operations Across Italian Basketball with ScorePlay

Lega Basket Serie A (LBA), the governing body for Italy's premier basketball...

04/03/2026

To Mark Five Years at Wrexham AFC, Co-Chairmen Rob Mac and Ryan Reynolds Will Host Broadcast During Wrexham-Swansea City on March 13

Wrexham AFC co-chairmen Rob Mac and Ryan Reynolds will host a first-of-its-kind ...

04/03/2026

FOX Sports Marks 100 Days to FIFA World Cup 2026 with Company-Wide Celebration

The countdown is underway and with just 100 days to go until the world's greatest sporting event begins on Thurs., June 11, FOX Sports, America's Englis...

04/03/2026

Exchange, NBCUniversal Team Up to Provide Service Members with Free Streaming of Paralympic Winter Games

No matter where they are in the world, service members and veterans can stream N...

04/03/2026

Telemundo Releases Somos Ms, the Official Anthem of its FIFA World Cup 2026 Coverage

Telemundo officially releases Somos M s, the anthem for the network's cove...

04/03/2026

Hollywood Professional Association Concludes 2026 HPA Tech Retreat

The Hollywood Professional Association (HPA) concluded the 2026 HPA Tech Retreat, convening more than 800 industry leaders, technologists, creatives, and execut...

04/03/2026

Case Study: How SEG+ Unified Utah's Biggest Sports Teams into a 40%+ Subscriber Growth Streaming Platform

Smith Entertainment Group transformed how local sports are consumed by creating ...

04/03/2026

SVG in Indy: Pacers Sports & Entertainment Remotely Produces Broadcasts of G League's Noblesville Boom

The production method, which spans a distance of 36.2 miles, was designed and im...

04/03/2026

Haivision Releases Seventh Annual Broadcast Transformation Report, Highlights Key Trends Shaping Live Production in 2026

Haivision, a global provider of mission-critical, real-time video networking and...

04/03/2026

SVG Sit-Down: Quantum CEO Hugues Meyrath on Reshaping the Company, the Impact of AI, Evolving Media Storage

When Hugues Meyrath came out of retirement to take the helm as CEO of Quantum, i...

04/03/2026

As Banana Ball Expands Exponentially, So Too Do Its Production Capabilities

Last week's launch of Banana Ball Championship League has spurred a significant upgrade of production facilities The Savannah Bananas, arguably the hottest...

04/03/2026

Press Release TEST

sldkfjsdlfkjsldkfjsldkjfslkdjfslkdjfsl Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus...

04/03/2026

APTS Announces Public Broadcast Leadership, Advocacy Awards

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

NBC Sports, USA Sports Extend Rights Deal with PGA

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

Broadcasters Gather in DC for NAB State Leadership Conference

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

Bloodhounds' Season 2 Gears Up for April 3 Premiere with Hard-Hitting New Teaser and Poster

Back to All News Bloodhounds' Season 2 Gears Up for April 3 Premiere with ...

04/03/2026

Netflix Ads Suite Expands Capabilities

Back to All News Netflix Ads Suite Expands Capabilities Business 04 March 2026 GlobalUnited States Link copied to clipboard After launching the Netflix Ad...

04/03/2026

March 02, 2026

Scripps Research welcomes healthcare innovator Joe Kiani to the Board of Directors Kiani brings decades of experience in patient safety and public service. Mar...

04/03/2026

March 03, 2026

Nanoparticle vaccine approach takes on a new target: Hepatitis C virus Scripps Research scientists reengineer critical proteins on the surface of HCV, paving th...

03/03/2026

LIV Golf, Beyond Sports Elevate Online Gaming Ecosystem with Launch of LIV Golf Fantasy and LIV X

Beyond Sports, a Sony group company, and LIV Golf, the world's golf league, ...

03/03/2026

Ilitch Sports + Entertainment Announces Launch of Detroit SportsNet

Ilitch Sports + Entertainment announces the launch of Detroit SportsNet (DSN), a year-round broadcast home for two of Detroit's franchises. With flexible op...

03/03/2026

Advanced Systems Group Promotes Gretchen Taipale to Vice President, Managed Services

Advanced Systems Group, LLC (ASG), a technology and services provider for media ...

03/03/2026

PGA of America, NBC Sports, and USA Sports Extend Media Rights Agreement Through 2033

The PGA of America, NBC Sports and USA Sports extend their media rights agreemen...

03/03/2026

HONOR, ARRI Announce Technical Collaboration to Bring ARRI Image Science into Next-Gen Consumer Devices

AI device ecosystem company HONOR enters into a strategic technical collaboratio...

03/03/2026

Telos Alliance Partners with College Radio Foundation to Support College Broadcasters

Cleveland's Telos Alliance, pioneers in broadcast technology for 30 years, l...

03/03/2026

Sennheiser Relaunches MD 9235 Wireless Mic Head

The MD 9235 microphone head for wireless handhelds has been a firm favorite with many engineers and artists for its ability to cut through high on-stage levels ...

03/03/2026

Haivision to Showcase Private 5G and Live Video Contribution Innovations at MWC 2026

Haivision Systems Inc. (Haivision), a global provider of mission-critical, real-...

03/03/2026

BMG Expands Washington Broadcast Center with 3 New TV Studios and Podcast Studio for Media Clients

Broadcast Management Group (BMG) announces the expansion of its 62,000-square-fo...

03/03/2026

Closing the Loop: Maroon 5 and the End of the Analog Era

Maroon 5's musical tour in 2025 marked a leap forward in live audio as Monitor Engineer Dave Rupsch utilized Sennheiser's all-digital Spectera wireless ...

03/03/2026

SVG in Indy: Pacers Sports & Entertainment Finds Production Sweet Spot in ST 2110-Based Control Center

Designed specifically for pro basketball, the renovated space at Gainbridge Fiel...

03/03/2026

Lawo Appoints Jamie Dunn CEO

As part of the move, former CEO Phillipp Lawo joins the broadcast-tech provider's Supervisory Board Lawo has announced appointment of Jamie Dunn as chief e...

03/03/2026

NBC Turns Back the Clock to 1990s for NBA Coast 2 Coast' Tuesday

A team of legendary announcers and analysts and a classic graphics look will bring the past to life NBC Sports and Peacock will return to yesteryear for tonigh...

03/03/2026

Sundance Film Festival: CDMX 2026 Returns for Its Third Edition

From April 30 to May 3, Sundance Film Festival: CDMX 2026 will offer a selection of exciting independent cinema. Mexico City, March 3, 2026 - At a moment of he...

03/03/2026

Magellan AI Integrates Nielsen DMA Data to Bring Local Market Measurement to Podcast Attribution

Nielsen's DMA data gives Magellan AI users a standardized way to measure th...

03/03/2026

Lawo Promotes Jamie Dunn to CEO

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Elements To Showcase Newly Unveiled GRID NAS Platform At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Moments Lab To Feature Agentic AI For Video Workflows At 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Marshall Electronics Launches Compact CV356-10X Full HD C...

Marshall Electronics premieres the CV356-10X, its latest compact 10X camera that offers Full HD with simultaneous SDI and HDMI outputs, at NAB 2026 (Booth C8339...

03/03/2026

farmerswife and Cirkus to Showcase Smarter Media Workflow...

farmerswife, the industry-leading enterprise operations platform for broadcast and post-production, today announced it will exhibit at NAB Show 2026 in Las Vega...

03/03/2026

Manfrotto ONE Hybrid Tripod Wins iF Design Award 2026

Manfrotto has announced that the Manfrotto ONE Hybrid tripod has won the iF DESIGN AWARD 2026, one of the world's most respected design honours. Selected ...

03/03/2026

DHD to Introduce Latest Generation Broadcast Audio Mixers...

DHD is expanding the capabilities of its DX2, RX2, SX2 and TX2 broadcast audio mixers, RM1 portable production unit and XC3/XD3/XS2 processing cores with the in...

03/03/2026

Synamedia and MoMe launch first streaming CDN in Spain

Leading video software provider Synamedia and MoMe, a leading Spanish consultancy and systems integrator, today announced the launch of Spain's first stream...

03/03/2026

Long-Awaited ATSC 3.0 Rulemaking Overshadows NAB Show Expectations

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Audio Tech at NAB Show: Are We in the Second Wave' of IP?

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

IP's Impact on Imaging Tech on Full Display at NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Live Production Over IP in 2026: Software-Defined Everything

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

NAB Show Leverages Revitalized LVCC To Reflect M&E Transformation

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

Home Post Production Strengthens Factual and Natural Hist...

Home Post Production has further expanded its factual, unscripted, and entertainment capabilities with the acquisition of Picture Shop Bristol, a leading post h...

03/03/2026

Iyuno Taps Dante AV to Sync Audio and Video Content

Share Copy link Facebook X Linkedin Bluesky Email...

03/03/2026

HBO Max and Paramount+ Streamers to Merge

Share Copy link Facebook X Linkedin Bluesky Email...