
Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.
NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.
AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.
But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.
Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.
Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.
NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:
NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.
NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.
NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.
Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.
With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.
Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:
Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service
Google Cloud's Vertex AI, Google Kubernetes Engine
Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service
Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine
Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.
For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.
Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.
The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.
From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.
Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.
Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.
Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
21/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
21/05/2026
For cinematographer Ashley Barron ACS ( Rivals , Dangerous Liaisons , Disney's Doctor Who ), setting the look of Netflix's How to Get to Heaven from ...
21/05/2026
Jean-Fran ois Hensgens, AFC, SBC, (Tueurs, Six Days in Spring, Le Fil) brings his signature visual approach to Cannes 2026 premiering biopic L'Affaire Marie...
21/05/2026
2026 delegation for Partner with Australia (UK) initiative announced 21 May 2026
Screen Australia and Ausfilm have today announced the delegation for the 2026 ...
21/05/2026
Beeble launches Canvas, a node-based AI compositor for VFX and Virtual Productio...
21/05/2026
Shelly Johnson Elected President of American Society of Cinematographers
Brie Clayton May 20, 2026
0 Comments
The American Society of Cinematographers...
21/05/2026
ARRI expands Management Board to accelerate next phase of growth and innovation
Brie Clayton May 20, 2026
0 Comments
ARRI expands its Management Board...
20/05/2026
True to the Music City theme, the Tennessee Titans plan to make live music a prominent fixture on game days at the new venue
Reflecting the continuing converge...
20/05/2026
The long-term goal is innovation for the next generation of hockey rinks at both the professional and the community level
The National Hockey League (NHL) toda...
20/05/2026
End-users and vendors alike work toward cost-effective yet high-quality solution...
20/05/2026
The National Football League has announced that Nashville will host Super Bowl LXIV in 2030 at the new Nissan Stadium. The announcement was made at the NFL Spri...
20/05/2026
The NFL has announced that the 2028 NFL Draft presented by Bud Light will take place in Minnesota, uniting fans from around the world to celebrate one of the mo...
20/05/2026
As professional AV workflows continue to evolve, expectations around audio, video, networking, and control are rising rapidly. At InfoComm 2026 in Las Vegas, La...
20/05/2026
Haivision will showcase its latest innovations at InfoComm 2026, taking place fr...
20/05/2026
For the seventh year, Spotify is returning to CMA Fest with Spotify House, the festival's premiere destination for fans. We're taking over downtown Nash...
20/05/2026
Amp-simulation software expanded
Acustica Audio's latest release greatly expands on their amp-simulation platform, turning it into a complete amplifica...
20/05/2026
MainStage integration, Analog Lab improvements & more
Arturia have just announced the release of an update that brings an assortment of new features to thei...
20/05/2026
At the Annual General Meeting held on May 20, 2026, the shareholders of SGL Carb...
20/05/2026
From Gulkula to the nation: Yothu Yindi Foundation and NITV deepen national acce...
20/05/2026
SBS appoints David Fernandez as National Manager, Digital & TV Sales
20 May, 2026
Media releases
SBS has appointed David Fernandez as National Manager, Dig...
20/05/2026
Rohde & Schwarz and INFOZAHYST: A strategic alliance set to redefine modern defe...
20/05/2026
A Rotating Detonation Engine being hot fire tested at Purdue University's Zu...
20/05/2026
A U.S. Army VAMPIRE system, assigned to Bravo Battery, 1st Battalion, 51st Air Defense Artillery Regiment, 7th Infantry Division/Multi-Domain Command - Pacific ...
20/05/2026
Cable Captures Only Monthly Increase Among Viewing Categories in March, Earns it...
20/05/2026
April brought a symbolic decrease in the overall time spent in front of television screens. On average, Poles watched video content for 3 hours and 51 minutes a...
20/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/05/2026
The Royal Television Society Technology Centre today announces the launch of the RTS Technology Awards 2026, celebrating excellence, innovation and achievement ...
20/05/2026
In the heart of London's financial district, the new purpose-built Troubadour Canary Wharf Theatre invites audiences to experience Suzanne Collins' inte...
20/05/2026
LiveU, the leader in live IP-video solutions, today announced that production powerhouse BCC Live successfully deployed the new LU900Q intelligent production un...
20/05/2026
Nella Mente di Narciso Docuseries Uses Blackmagic Design Workflow
Brie Clayton May 19, 2026
0 Comments
PYXIS 6K full frame camera and DaVinci Resolve ...
20/05/2026
Beeble launches Canvas, a node-based AI compositor for VFX and Virtual Productio...
20/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/05/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
20/05/2026
First Nations Factual Co-production Development Fund launched to elevate Indigen...
20/05/2026
TLC Announces All-Star Comedy Line-Up for Mock the Week: Summer Specials' o...
20/05/2026
Back to All News
One Hundred Years of Solitude Concludes This August, With Part...
20/05/2026
Back to All News
Beloved NHK Dramas to Stream on Netflix Worldwide Starting June 22
Entertainment
20 May 2026
GlobalJapan
Link copied to clipboard
Netflix...
20/05/2026
May 20 2026, 06:00 (PDT) Dolby Recognized as 2025 Supplier of the Year and Over...
20/05/2026
Mayo's Dee Freney and Margaret Leahy from Galway have reached the final of RT Today's TV Home Cook competition.
Both contestants will cook again live...
20/05/2026
RT IN FULL BLOOM AT BORD BIA BLOOM 2026 WITH LIVE BROADCASTS, MUSIC, CHAT AND M...
19/05/2026
The winner of Thomson Foundation's Young Journalist of the Year 2025, Tracy Bonareri Onchoke, and runner up Wangu Kanuri enjoyed a three-day trip to London ...
19/05/2026
Cisco and the USGA have announced a multiyear extension of their partnership, which began in 2018. Cisco serves as the Official Technology Partner of the USGA, ...
19/05/2026
Urban Edge Network (UEN), a streaming platform for NAIA sports, has announced a partnership with Spiideo to provide streaming and production tools to UEN's ...
19/05/2026
Warner Bros. Discovery (WBD) will provide live coverage of all 900 Roland-Garros matches across its platforms beginning with qualifiers on May 18. In Europe, 21...
19/05/2026
Tubi, Fox Corporation's free streaming service, has announced the launch of the FIFA World Cup 2026 FOX Hub, a dedicated destination for World Cup programmi...