
Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.
NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.
AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.
But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.
Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.
Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.
NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:
NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.
NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.
NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.
Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.
With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.
Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:
Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service
Google Cloud's Vertex AI, Google Kubernetes Engine
Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service
Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine
Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.
For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.
Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.
The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.
From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.
Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.
Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.
Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
North America Stories
21/07/2026
With winners Spain having already paraded the World Cup trophy through Madrid in front of millions of fans and the dust now settling on an epic tournament of th...
21/07/2026
In a time of disarray in regional sports broadcasting, the league will provide c...
21/07/2026
When NEP designed one of its largest recurring mobile production unit fleets for PGA TOUR
coverage, the leading media services provider faced a familiar challe...
21/07/2026
Warner Bros. Discovery's unbroken lead is interrupted as sports and live mus...
21/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/07/2026
Across the TV and streaming ecosystem, industry players face mounting pressure to do more with less. They must simplify operations, protect premium content and ...
21/07/2026
Tuxera, a leading prover of quality-assured file systems and networking technologies, is bringing its latest advances in connectivity performance to the media w...
21/07/2026
COW Jobs: Apple Motion Expert for Custom FCP Rigged Graphics Package
Brie Clayton July 21, 2026
0 Comments
HIRING: Apple Motion Expert for Custom FCP ...
21/07/2026
Cinegy GmbH, the premier provider of software-defined television technology, will use its presence at IBC2026 (stand 7.A01, Amsterdam RAI, 11 14 September) to...
21/07/2026
Calrec will be located in Hall 8, on Stand C47
Beyond bigger
For years, broadcast facilities were built around one assumption: provision for the biggest pro...
21/07/2026
Pebble, the leading automation, content management and integrated channel specialist, will discuss its future-facing developments for the new generation of medi...
21/07/2026
nxtedition returns to IBC2026 with its consolidated production platform, showing how scripting, editing, graphics, AI-assisted tools and automation can work wit...
21/07/2026
C2PA content signing helps broadcasters, public institutions, and publishers give audiences a verifiable record of where their video content came from and how i...
21/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/07/2026
New capability centralizes collaboration, feedback, security and status tracking within VIDA, providing a structured, auditable approach to master quality contr...
21/07/2026
Lightcraft Launches Exclusive Spark Story Beta for Filmmakers and Creators at SI...
21/07/2026
Indie Road Movie Where in the Hell Shot with Pocket Cinema Camera 4K
Brie Clayton July 20, 2026
0 Comments
Colorist blends vintage film looks to shape...
21/07/2026
Which USS Defiant Pulse Phaser effect is better?
Graham Quince July 20, 2026
0 Comments
Aargh, ever since @DarkRavenProductions posted a comment ask...
21/07/2026
Every live event starts with a defining moment. A camera captures the game-winni...
21/07/2026
AI has entered the gigascale era.
The world's most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier mod...
21/07/2026
NVIDIA Vera Rubin is here, and it's going gigascale.
Vera Rubin NVL72 produ...
20/07/2026
The former Northwestern broadcast-operations leader is helping train the next generation of live-sports-production talent across the Big Ten
The sports-product...
20/07/2026
By Lucy Spicer
One of the most exciting things about the Sundance Film Festival is having a front-row seat for the bright future of independent filmmaking. Whi...
20/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/07/2026
Starfish Technologies will use IBC2026 to showcase the flexibility of its transport stream processing software, including the latest versions of TS Splicer (Win...
20/07/2026
Bitfocus, the specialist in media control and monitoring, will show at IBC2026 (Elgato stand 8.D31, Amsterdam RAI, 11 14 September) how its Buttons control la...
20/07/2026
Mediagenix, a global leader in smart content solutions to profitably connect the right content to the right audience, today announced new AI capabilities that e...
20/07/2026
Big Blue Marble's Cloud DRM has been nominated in the DRM/Content Protection category of the 2026 Streaming Media Readers' Choice Awards.
Only four pro...
20/07/2026
At this year's SIGGRAPH conference, running through Thursday, July 23, in Lo...
20/07/2026
Erin Davis calls it the SuperDuperPOD. That's two things in one name: phar...
19/07/2026
Justin Bieber, Madonna, Shakira, BTS make for a diverse lineup, and the venue ad...
19/07/2026
The stage is set: three-time champion Argentina will defend its World Cup title ...
18/07/2026
Topics include pre-match ceremonies, live performances, the tournament's fir...
18/07/2026
When FIFA and HBS set out to produce the 2026 FIFA World Cup, the numbers alone ...
18/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
18/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
17/07/2026
In-venue and creative video staffers at the professional and collegiate level ha...
17/07/2026
Production workers at Brooklyn Bowl's Williamsburg location voted 15-1 to join IATSE Local 4. The bargaining unit covers 24 production workers at the venue,...
17/07/2026
DAZN and ADI Predictstreet have announced an exclusive global strategic partners...
17/07/2026
Zixi and Comcast Technology Solutions (CTS) have announced a strategic integrati...
17/07/2026
Professional Fighters League (PFL) has announced a multi-year partnership with E...
17/07/2026
Spectrum Business has announced Spectrum TV Control Pro, a centralized app-based...
17/07/2026
Clark Wire and Cable has announced that Rick Fernandez, Managing Director of Axxion Consulting, will serve as Independent Manufacturers Representative for Centr...
17/07/2026
TikTok, the NBA, and the WNBA have announced a multi-year global content partnership covering highlights distribution, creator access to marquee events, live-ga...
17/07/2026
Company alleges article contained false and misleading claims regarding customer data...
17/07/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...