
Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.
NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.
AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.
But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.
Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.
Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.
NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:
NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.
NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.
NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.
Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.
With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.
Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:
Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service
Google Cloud's Vertex AI, Google Kubernetes Engine
Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service
Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine
Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.
For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.
Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.
The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.
From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.
Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.
Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.
Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
Most recent headlines
11/12/2025
Dalet, a leading provider of cloud-native, end-to-end media workflow solutions, ...
18/11/2025
Independent media across Central America are operating under intensifying financial pressure, yet there is a clear appetite for models that can sustain both ind...
18/11/2025
NBA Debuts New Comms Infrastructure and Systems for RefereesTwo-phase rollout is intended to improve game flow, enhance officiating accuracyBy Dan Daley, Audio ...
18/11/2025
Platinum White Paper: How Aggreko Delivers Certainty For Broadcasters On The Wor...
18/11/2025
Kiswe Extends DTC Products With Kiswe Core Cloud-Based Tool for Distributing Con...
18/11/2025
Spanish Basketball Federation partners with ScorePlay to power digital transform...
18/11/2025
HBS selects BBright encoders and decoders to connect ST 2110 live production wor...
18/11/2025
Ashes to Ashes: Inside TNT Sports' hybrid and maverick' production plan...
18/11/2025
SVG All-Stars: Mimi Fotopoulos, Director, Talent and Production Operations, Tenn...
18/11/2025
SVG LIVE! Conference Explores the Tech Side of Sports and Entertainment ConvergenceLive sports and music productions increasingly share gear and infrastructureB...
18/11/2025
LPGA Ups Its Production Game in 2026 With 50% More Cameras, SSMO's & Drones,...
18/11/2025
ESPN's Dunk the Halls Real-Time Animated NBA Game Set to Return for Christma...
18/11/2025
By Jessica Herndon
When Chilean filmmakers Diego C spedes and Giancarlo Nasi ar...
18/11/2025
Earlier this year, we announced the BAPE x SPOTIFY x SYNA by Central Cee (aka Cench) capsule collection, a collaboration that blends sound, style, and street c...
18/11/2025
Our mission to make Spotify the ultimate home for all things audio continues. St...
18/11/2025
SGL Carbon and the renowned Link ping University inaugurated an advanced coating...
18/11/2025
CANBERRA, Australia, Nov. 18, 2025 - L3Harris Technologies (NYSE: LHX) has launched a new, next-generation software device called NETCASTER - the NETwork Contr...
18/11/2025
Pictured are L3Harris ISR President Jason Lambert and Waleid Al Mesmari, EDGE President-Space & Cyber Technologies, signing with Carlo Igniades, L3Harris Region...
18/11/2025
During October, streaming's share of TV viewing in Mexico settled at 23.7%, a marginal shift of -0.8 share points from the previous month.
Disclaimer: YUMI...
18/11/2025
Broadcast Builds On Lead Over Cable, Driven by Football and Drama Programming Ga...
18/11/2025
CINCINNATI E.W. Scripps has issued a statement responding to news that Sinclair has acquired approximately 8.2% of the outstanding class A (non-voting) shares o...
18/11/2025
TYSONS, Va. Tegna has announced that its shareholders have voted overwhelmingly to approve the proposed $6.2 billion merger with Nexstar Media....
18/11/2025
NEW YORK Short-form vertical video has exploded across platforms like TikTok, Instagram, and YouTube, but a new survey commissioned by Media.net, a provider of ...
18/11/2025
SHENZHEN, China DJI today unveiled the Osmo Action 6, an all-in-one action camera featuring a variable aperture with a range from f/2.0 to f/4.0, the company...
18/11/2025
In the high-stakes world of film sound, there's no room for second chances. When an actor whispers a line with the weight of an entire scene, or lets loose ...
18/11/2025
Sonnet Technologies is having a sale on a selection of popular products, including a Thunderbolt 4 dock, an eGPU chassis, SATA and M.2 SSD PCIe cards, and mor...
18/11/2025
Amagi, a cloud-based SaaS technology solutions provider for broadcast and streaming TV, today announced the launch of 4Fangs, a new Free Ad-Supported Streaming ...
18/11/2025
Women In Media (WiM) announces its honorees for the 2025 Holiday Toast. The annual celebration recognizes legendary creatives whose work uplifts and inspires th...
18/11/2025
Cinematographer Adam Newport-Berra ( Good Fortune , The Bear , Euphoria ) emerged from the 2025 Emmy season with a statuette celebrating his Outstanding Cinem...
18/11/2025
Atlanta-based gaffer and lighting programmer Quinton Thomas brings the same practical, problem-solving instinct of a board op to every set he walks onto. With a...
18/11/2025
Suitelife Systems, a division of NFB Consulting Group, headquartered in California, has announced the appointments of Charles Sotto as Director of Media Technol...
18/11/2025
Sound Director and Audio Product Manager Scott Kramer has built a career around shaping stories through audio, guiding projects across film, television and stre...
18/11/2025
iWedia, a global leader in software solutions for connected TV devices, announces its participation at the APAC TV Summit 2025, taking place November 18 20 in B...
18/11/2025
Mark Roberts Motion Control (MRMC) is proud to announce the launch of Flair Bridge, a groundbreaking solution that redefines how operators control MRMC's in...
18/11/2025
Over the summer, world famous musician and rapper Drake continued the rollout of his new album Iceman, by live streaming the second and third episodes of a seri...
18/11/2025
KEIZER, Ore. Municipal broadcaster Keizer City Television, K23 TV, has deployed four Telycam PTZ cameras to upgrade the visual quality of its live meeting cover...
18/11/2025
STAMFORD, Conn. NBC Sports and Peacock have announced that they are working for the second consecutive year with the National Football League, EA Sports and Gen...
18/11/2025
LONDON A new Ampere Analysis study finds that familiar franchises are successfully driving kids' TV consumption on Netflix and that the streamers big bet on...
18/11/2025
PORTSMOUTH, N.H. TV viewers continue to find it challenging to find relevant programs in the fragmented universe of streaming content, according to new findings...
18/11/2025
20 Bob Dylan Songs That Reflect a Legacy of A-Changin On the heels of Bob Dylan receiving a Berklee honorary doctorate, we take stock of one of the most singu...
18/11/2025
Abu Dhabi, UAE, November 18, 2025 At Dubai Airshow 2025, Mira Aerospace, the H...
18/11/2025
Rohde & Schwarz and MILTON expand partnership, unveiling new RF spectrum monitor...
18/11/2025
Rohde & Schwarz presents multi-purpose R&S NGT3600 high-precision dual-channel p...
18/11/2025
Rohde & Schwarz and ELT Group: Towards a Global Strategic Partnership Across All...
18/11/2025
Wuppertal November 18, 2025
Tisman Service Expands Rental Offerings with Riede...
18/11/2025
Today, Microsoft, NVIDIA and Anthropic announced new strategic partnerships. Anthropic is scaling its rapidly growing Claude AI model on Microsoft Azure, powere...
18/11/2025
AI agents have the potential to become indispensable tools for automating complex tasks. But bringing agents to production remains challenging.
According to Ga...
18/11/2025
DJ Carey: The Dodger starts Monday 24th November at 9.35pm on RT One and RT Pl...
17/11/2025
EA SPORTS Madden NFL Cast to Return Thanksgiving Night With Immersive, Data-Driv...
17/11/2025
Behind the Broadcast Booth: Impact Ventures' Greer Christian on Building Her...