Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

North America Stories

02/02/2026

Inside the Avionics That Make Artemis II Possible: How L3Harris Units Drive the Space Launch System

Photo Credit: NASA. Space Launch System (SLS) rocket and Orion Spacecraft rollou...

02/02/2026

ESPNs deal for NFL Network Gets Regulatory Approval

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

02/02/2026

Hewshott Entering New Era as Daniel Lee Transitions to G...

Hewshott, an industry leading global AV, IT, Theatre, and Acoustics consultancy firm has completed a global transition with current UK Managing Director, Daniel...

02/02/2026

Public Media Management Names LTN as Technology Partner f...

Public Media Management (PMM) today announced LTN as the technology partner for PMM Cloud, its new managed, cloud-based master control solution purpose-built fo...

02/02/2026

Big Academic Investments in Virtual Production Fuel Adoption

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

02/02/2026

Audiences Can Expect Seamless Viewing From Milan-Cortina

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

02/02/2026

NBC Olympics Will Take Audio to New Heights in Milan-Cortina

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

02/02/2026

XR Acquires Telly Traffic to Bolster UK Advertising Opera...

XR, the leading platform powering advertising operations, today announced the acquisition of Telly Traffic, a UK-based business affairs specialist with nearly t...

02/02/2026

Big Blue Marble Becomes Launch Partner for the AWS Europe...

Big Blue Marble, a provider of broadcast-grade, cloud-native video solutions for broadcasters, service providers, and content owners, has become a launch partne...

02/02/2026

Cesc Gay's New Film Premieres March 27 on Netflix

Back to All News Cesc Gays New Film Premieres March 27 on Netflix Entertainment 02 February 2026 GlobalSpain Link copied to clipboard Download the first i...

31/01/2026

US Navy Selects L3Harris Red Wolf for Precision Attack Strike Munition Program

The Navy's Air Test and Evaluation Squadron (HX) 21 launch a Long Range Attack Missile from an AH-1Z off coast of Virginia in late 2025. This demonstration ...

31/01/2026

creativespace 3 0 5 Simplifies Creative Workflows with In...

DigitalGlue, creator of the award-winning creative.space Platform, has announced the release of creative.space OS 3.0.5, the latest software update within the ...

31/01/2026

ES Broadcast Hire sends out kit for major February sporti...

ES Broadcast Hire, the long-established hire arm of ES Media Group, has spent the last few months busily preparing and sending out high-quality equipment for a ...

31/01/2026

ISE: Sony Electronics Launches BRAVIA Professional Displays BZ-P Series

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

31/01/2026

Fox Sports Details FIFA World Cup 2026 Broadcast Schedule

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

31/01/2026

ISE: LYNX Technik Debuts New 4-Channel 12G-SDI Fiber Converters

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

31/01/2026

PBS Promotes Scott Nourse to Chief Technology Officer

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

2026 Sundance Film Festival Announces Award Winners

Top L-R: The Friend's House is Here, Josephine, The Lake, Bedford Park, Who Killed Alex Odeh? Second Row L-R: Take Me Home, American Pachuco: The Legend of...

30/01/2026

Artemis II Wet Dress Rehearsal: Taking Things Into the Home Stretch

The Artemis II wet dress rehearsal will simulate the launch countdown, fully loading fuel and verifying systems ahead of the first SLS and Orion crewed flight....

30/01/2026

Copper Leaf Media Merges With Dimension PR

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

MXL: Aligning Broadcast Production With Software

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

Grass Valley and NETGEAR Partner to Accelerate Enterprise...

Grass Valley , the leading technology provider for live production solutions, and NETGEAR Inc. (NASDAQ: NTGR), a global leader in network solutions, today anno...

30/01/2026

tvONE Appoints Amit Singh as Regional Sales Manager for I...

tvONE, a leading video processor, signal distribution technology and media server developer, announces the expansion of Amit Singh's role to Regional Sales ...

30/01/2026

Mike Aiton Brings Stories to Life with NUGEN Audio

With a career that spans four decades across television, film and post-production, Freelance Sound Designer and Post-production Sound Mixer Mike Aiton has built...

30/01/2026

DPA Showcases its Complete Wireless Microphone Ecosystem...

DPA Microphones will feature its new, fully integrated wireless microphone ecosystem, designed to let audio professionals work faster, cleaner and with total co...

30/01/2026

Ventum Tech and Emergent Announce a Strategic Partnership...

As the Middle East continues to accelerate investment in next-generation media, broadcast, and immersive content technologies, Ventum Tech today announced a str...

30/01/2026

MRMC Broadcast to Highlight High-Precision Motion Control...

Mark Roberts Motion Control (MRMC), a Nikon company and global leader in robotic camera systems, today announced its participation at Integrated Systems Europe ...

30/01/2026

Peacock Hits 44 Million Subs, Lost $552 Million in Q4

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

ATSC Board Leadership Re-Elected For 2026

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

MRC Issues Final Digital Advertising Auction Transparency Standards

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

CBS Atlanta Launches New Weekday Morning News Show with AR/VR Set

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

30/01/2026

Boston Conservatory at Berklee Hosts the National Opera Association's 2026 Conference

Boston Conservatory at Berklee Hosts the National Opera Association's 2026 C...

30/01/2026

Student Spotlight: Sriram Narayanan

Student Spotlight: Sriram Narayanan The classical pianist shares his experience growing up with a language disability and finding his voice through music. Ja...

30/01/2026

2026 Media Industry Trends to Watch: Change is Your Competitive Advantage

Heading into 2026, the pace of change across radio, TV, and digital media is reaching an inflection point. Audience behaviors continue to evolve, measurement mo...

30/01/2026

The Danish Crime Series The Asset' Returns for a Second Season

Back to All News The Danish Crime Series The Asset' Returns for a Second Season Entertainment 30 January 2026 GlobalDenmark Link copied to clipboard ...

29/01/2026

L3Harris Technologies Reports Strong Full Year and Fourth Quarter 2025 Results, Initiates 2026 Guidance

MELBOURNE, Fla., January 29, 2026 - L3Harris Technologies (NYSE: LHX) reports fu...

29/01/2026

Nielsen Announces 2025 ARTEY Award Winners Following Record-Breaking Year of Streaming

Bluey' Wins Second Consecutive Top Streaming Title of the Year with 45 Billi...

29/01/2026

Report: Performance TV Ties With Social Media in Driving Ad Results

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

29/01/2026

ISE: NDI and OBSBOT Expand Partnership

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

29/01/2026

NTCA Asks FCC to Block Nexstar, Tegna Deal

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

29/01/2026

FCC Announces Tentative Agenda for February Open Meeting

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

29/01/2026

CBS Sports AFC Championship Game Attracts 48.6 Million Viewers

Share Share by: Copy link Facebook X Linkedin Bluesky Email...

29/01/2026

Boston Conservatory Orchestra Presents East Coast Premiere of Peter and Leonardo Dugan Piano Concerto

Boston Conservatory Orchestra Presents East Coast Premiere of Peter and Leonardo...

29/01/2026

Mercedes-Benz Unveils New S-Class Built on NVIDIA DRIVE AV, Which Enables an L4-Ready Architecture

Mercedes-Benz is marking 140 years of automotive innovation with a new S-Class b...

29/01/2026

'Love is Blind: Sweden' Returns for a Third Season - Premiering on March 12

Back to All News Love is Blind: Sweden Returns for a Third Season - Premiering ...

29/01/2026

Unmask Bridgerton' Season 4 With Our Complete Coverage Guide

Back to All News Unmask Bridgerton' Season 4 With Our Complete Coverage Guide Yerin Ha as Sophie Baek and Luke Thompson as Benedict Bridgerton in Season ...

29/01/2026

Extraordinary Crime Mysteries, Mythical Worlds and High-Stakes Psychological Thrillers: Inside Netflix's 2026 Chinese-Language Slate

Back to All News Extraordinary Crime Mysteries, Mythical Worlds and High-Stakes...

29/01/2026

Into the Omniverse: Physical AI Open Models and Frameworks Advance Robots and Autonomous Systems

Editor's note: This post is part of Into the Omniverse, a series focused on ...

29/01/2026

GeForce NOW Brings GeForce RTX Gaming to Linux PCs

Get ready to game - the native GeForce NOW app for Linux PCs is now available in beta, letting Linux desktops tap directly into GeForce RTX performance from the...

28/01/2026

2026 Sundance Film Festival Reveals Short Film Program Award Winners

Top L-R: The Liars, Jazz Infernal, Living with a Visionary Second Row L-R: Paper Trail, The Baddest Speechwriter of All, Crisis Actor Third Row: The Boys and ...