Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

North America Stories

05/06/2026

LSUs Sarah Ramundt on Celebrating the One-Year Anniversary of The Brand

The university-wide initiative has pushed their creative content to a new level Although collegiate athletics at a single institution can contain numerous spo...

05/06/2026

SVG GameDay, Ep. 18: New York Islanders Brooklyn Boyars - A Career Filled With Championships

In-venue and creative video staffers at the professional and collegiate level ha...

05/06/2026

Synamedia Enters New Chapter Following Lumine Group Acquisition Agreement

Synamedia has announced that Lumine Group has agreed to acquire its Video Network business. The company is positioning the transition as the start of a new phas...

05/06/2026

Ateme, BCE, and Scaleway Announce Sovereign Cloud Media Supply Chain Partnership

Ateme, Broadcasting Center Europe (BCE), and Scaleway have announced a strategic partnership to deliver a cloud-based media supply chain covering ingest through...

05/06/2026

Audinate Adds Three Models to Dante AVIO Install Adapter Line

Audinate has announced three new additions to its Dante AVIO Install adapter series: a 4-Channel Analog Input, a 4-Channel Analog Output, and a 2-Ch In/2-Ch Out...

05/06/2026

Sportradar Extends Data and AV Betting Rights Agreement for Wimbledon Beyond 2026

Sportradar Group AG has announced a multi-year extension of its exclusive global...

05/06/2026

M6 Group to Broadcast FIFA World Cup 2026 in 4K UHD via Fransat Satellite Platform

The M6 Group will broadcast FIFA World Cup 2026 matches live and in Ultra High D...

05/06/2026

The Athletic Adds PGA TOUR Highlights to Golf Coverage

The Athletic has announced that PGA TOUR highlights will be integrated into its golf coverage beginning with the Memorial Tournament presented by Workday. PGA T...

05/06/2026

Op-Ed: Beyond the Stadium: Why Real-Time Visual Intelligence Will Define World Cup 2026 Security

When the FIFA World Cup arrives in North America in 2026, it will bring more tha...

05/06/2026

FIFA+ Launches Exclusively on DAZN

FIFA and DAZN have announced the launch of FIFA exclusively on DAZN, consolidating FIFA's content portfolio within DAZN's sports platform. The move fol...

05/06/2026

Telemundo Announces Group Stage Broadcast Schedule and Talent Assignments for FIFA World Cup 2026

Telemundo's exclusive Spanish-language coverage of the FIFA World Cup 2026 G...

05/06/2026

Peacock to Stream Telemundos FIFA World Cup 2026 Coverage in Dolby Vision and Dolby Atmos

Dolby Laboratories and NBCUniversal have announced that Peacock will stream Tele...

05/06/2026

Formula 1 Extends Las Vegas Grand Prix Through 2037

Formula 1 has announced a 10-year extension to keep the Las Vegas Grand Prix on the F1 calendar through 2037. Las Vegas Grand Prix, Inc., Clark County, and the ...

05/06/2026

Inside SailGP New York: How Riedel Provides Tech Backbone for Challenging Production Effort

The broadcast-engineering team overcomes wind, speed, and salt water - and dista...

05/06/2026

Virtual Eye Powers Broadcast Graphics Across Golf Productions in Busy Three-Event Week

Deploying both onsite and remote crews, the company is providing calibrated-came...

05/06/2026

ESPN Pivots From REMI to Full On-Site Production for UFL Playoff Game on Daytonas Blank Canvas

With Inter&Co Stadium unavailable, ESPN's UFL team rebuilt its broadcast pla...

05/06/2026

Free Registration for SVG Regional Sports Production Summit Ends Today!

One of the most exciting and informative events on the SVG annual event calendar is the Regional Sports Production Summit, an annual gathering of industry profe...

05/06/2026

FOX Returns to Saratoga Race Course for 158th Running of Belmont Stakes

For the race's third year at the historic racetrack, the broadcaster has added cameras and will incorporate multiple drones The 158th edition of the Belmon...

05/06/2026

Ratings Roundup: NBCs Spurs-Thunder Game 7 Best Since 2016, ESPNs Stanley Cup Final Opener Hits Seven-Year High

Ratings Roundup is a rundown of recent rating news and is derived from press rel...

05/06/2026

Nielsen, Mediaocean Announce New Integration to Power The Next Phase of Data-Driven Linear, Advanced Audience Measurement

MRI-Simmons and S&P Global Mobility are expanding advanced audience capabilities...

05/06/2026

Latest Nielsen data reveals insurance advertising climbs as Australians weigh cost, cover and loyalty

New Nielsen data shows insurance ad spend grew 11%, while consumers remain highl...

05/06/2026

Nielsen Adds New Audience Data Partnerships

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

ASG Promotes Joe Marchitto to Western Regional CTO

ASG Promotes Joe Marchitto to Western Regional CTO Brie Clayton June 5, 2026 0 Comments Appointment to Support Engineering Alignment and Client Experi...

05/06/2026

Stargate Studios Colombia Uses DaVinci Resolve Studio for Vertical Microdramas

Stargate Studios Colombia Uses DaVinci Resolve Studio for Vertical Microdramas Brie Clayton June 5, 2026 0 Comments End to end post in one platform al...

05/06/2026

People Need to Come First When We Use AI

People Need to Come First When We Use AI Andy Marken June 5, 2026 0 Comments It's just surviving. Life's very existence requires destruction....

05/06/2026

Fahad Haider Joins NESN as VP, Operations and Engineering

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

GatesAir Opens New Brazil Office to Back DTV+ Rollout

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

Supreme Court Upholds FCC's Authority to Levy Fines

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

Montclair State University Will Run NJ PBS

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

Merkhet's Sam Matheny Urges Congress to Expedite BPS Deployment

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

Sinclair Launches NextGen TV Campaign in Columbus, Ohio

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

NAB Releases New Keep the Game On' Spot

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

NCTC, ACAC Going to Disney World for Independent Show

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

Nielsen Announces New Integrations for Improved Measurement

Share Copy link Facebook X Linkedin Bluesky Email...

05/06/2026

Frequency Launches In-Scene Advertising to Accelerate Str...

Frequency, the engine powering many of the world's leading streaming television channels, today announced the launch of In-Scene Advertising, a new monetiza...

05/06/2026

Berklee Study Reveals Video Has Become Essential to Music Careers

Berklee Study Reveals Video Has Become Essential to Music Careers Survey findings show social platforms have become the primary source of music for video cont...

04/06/2026

Sony's New PTZ Cameras Deliver 4K 60p; New STARVIS Sensor Meets Low-Light Demands

Sony Electronics is introducing the SRG-AS10, a 4K 60p-compatible PTZ auto-frami...

04/06/2026

SVG Students To Watch: Alex Albert, Texas A&M University

This recent grad from Spring, TX, led creative-video output for the Aggies' men's basketball team last season and has been producing video and creating ...

04/06/2026

USGA Brings AI Recaps, 3D Range Tracking, and Predictive Shot Tracing to U.S. Womens Open

For the first time at a women's golf major, every player in the field will r...

04/06/2026

Panasonic PT-RQ45 Projectors Power Lille Video Mapping Festival Opera House Installation

Three Panasonic PT-RQ45 40,000-lumen 3-Chip DLP projectors made their first live...

04/06/2026

Bitmovin and Akamai Support NRJ Groups Deployment of Akamai Adaptive Media Player 2

Bitmovin and Akamai have announced a collaboration with NRJ Group, a French mult...

04/06/2026

Telestream to Exhibit at InfoComm 2026 with Live Production and Media Workflow Demonstrations

Telestream will exhibit at InfoComm 2026 (Booth N7952), demonstrating media work...

04/06/2026

Sony Announces RIALTO 65 Image Sensor Block for VENICE 2, Targeting 2027 Release

Sony has announced the development of RIALTO 65, a 65mm format image sensor block for the VENICE 2 digital cinema camera, targeting release in the first half of...

04/06/2026

KOKUSAI DENKI Electric America to Exhibit 4K and Remote Production Solutions at InfoComm 2026

KOKUSAI DENKI Electric America will exhibit at InfoComm 2026 (Booth N8025, June ...

04/06/2026

Bell Media to Carry All 104 FIFA World Cup 2026 Matches Across TSN, RDS, and Streaming Platforms

Bell Media's TSN and RDS are the exclusive Canadian broadcasters of FIFA Wor...

04/06/2026

MASV Case Study: How MASV Reduced Miami HEAT's Road Game Video Transfer Times by 85%

The Challenge: Receiving Heavy Media Files From Road Games Quickly and ReliablyT...

04/06/2026

MASV Outlines Seven-Step Sports Analytics Workflow, Highlights File Transfer as Key Bottleneck

MASV, a managed file transfer platform used in broadcast and live sports product...

04/06/2026

NESN Appoints Fahad Haider as Vice President of Operations and Engineering

NESN has announced the appointment of Fahad Haider as Vice President of Operations and Engineering. Haider returns to NESN, where he previously served as Vice P...

04/06/2026

Sports Broadcaster, Executive, and Author David J. Halberstam Dies

David J. Halberstam, who spent almost 50 years in sports as a broadcaster and an executive, died June 2 after a years-long battle with brain cancer. Over his l...

04/06/2026

Grass Valleys Ben Dolinky on Offering Teachable Technology to College Students Across the Country

Although collegiate production programs are tasked with delivering high-quality ...