Sony Pixel Power calrec Sony

Fast, Low-Cost Inference Offers Key to Profitable AI

23/01/2025

Businesses across every industry are rolling out AI services this year. For Microsoft, Oracle, Perplexity, Snap and hundreds of other leading companies, using the NVIDIA AI inference platform - a full stack comprising world-class silicon, systems and software - is the key to delivering high-throughput and low-latency inference and enabling great user experiences while lowering cost.

NVIDIA's advancements in inference software optimization and the NVIDIA Hopper platform are helping industries serve the latest generative AI models, delivering excellent user experiences while optimizing total cost of ownership. The Hopper platform also helps deliver up to 15x more energy efficiency for inference workloads compared to previous generations.

AI inference is notoriously difficult, as it requires many steps to strike the right balance between throughput and user experience.

But the underlying goal is simple: generate more tokens at a lower cost. Tokens represent words in a large language model (LLM) system - and with AI inference services typically charging for every million tokens generated, this goal offers the most visible return on AI investments and energy used per task.

Full-stack software optimization offers the key to improving AI inference performance and achieving this goal.

Cost-Effective User Throughput Businesses are often challenged with balancing the performance and costs of inference workloads. While some customers or use cases may work with an out-of-the-box or hosted model, others may require customization. NVIDIA technologies simplify model deployment while optimizing cost and performance for AI inference workloads. In addition, customers can experience flexibility and customizability with the models they choose to deploy.

NVIDIA NIM microservices, NVIDIA Triton Inference Server and the NVIDIA TensorRT library are among the inference solutions NVIDIA offers to suit users' needs:

NVIDIA NIM inference microservices are prepackaged and performance-optimized for rapidly deploying AI foundation models on any infrastructure - cloud, data centers, edge or workstations.

NVIDIA Triton Inference Server, one of the company's most popular open-source projects, allows users to package and serve any model regardless of the AI framework it was trained on.

NVIDIA TensorRT is a high-performance deep learning inference library that includes runtime and model optimizations to deliver low-latency and high-throughput inference for production applications.

Available in all major cloud marketplaces, the NVIDIA AI Enterprise software platform includes all these solutions and provides enterprise-grade support, stability, manageability and security.

With the framework-agnostic NVIDIA AI inference platform, companies save on productivity, development, and infrastructure and setup costs. Using NVIDIA technologies can also boost business revenue by helping companies avoid downtime and fraudulent transactions, increase e-commerce shopping conversion rates and generate new, AI-powered revenue streams.

Cloud-Based LLM Inference To ease LLM deployment, NVIDIA has collaborated closely with every major cloud service provider to ensure that the NVIDIA inference platform can be seamlessly deployed in the cloud with minimal or no code required. NVIDIA NIM is integrated with cloud-native services such as:

Amazon SageMaker AI, Amazon Bedrock Marketplace, Amazon Elastic Kubernetes Service

Google Cloud's Vertex AI, Google Kubernetes Engine

Microsoft Azure AI Foundry coming soon, Azure Kubernetes Service

Oracle Cloud Infrastructure's data science tools, Oracle Cloud Infrastructure Kubernetes Engine

Plus, for customized inference deployments, NVIDIA Triton Inference Server is deeply integrated into all major cloud service providers.

For example, using the OCI Data Science platform, deploying NVIDIA Triton is as simple as turning on a switch in the command line arguments during model deployment, which instantly launches an NVIDIA Triton inference endpoint.

Similarly, with Azure Machine Learning, users can deploy NVIDIA Triton either with no-code deployment through the Azure Machine Learning Studio or full-code deployment with Azure Machine Learning CLI. AWS provides one-click deployment for NVIDIA NIM from SageMaker Marketplace and Google Cloud provides a one-click deployment option on Google Kubernetes Engine (GKE). Google Cloud provides a one-click deployment option on Google Kubernetes Engine, while AWS offers NVIDIA Triton on its AWS Deep Learning containers.

The NVIDIA AI inference platform also uses popular communication methods for delivering AI predictions, automatically adjusting to accommodate the growing and changing needs of users within a cloud-based infrastructure.

From accelerating LLMs to enhancing creative workflows and transforming agreement management, NVIDIA's AI inference platform is driving real-world impact across industries. Learn how collaboration and innovation are enabling the organizations below to achieve new levels of efficiency and scalability.

Serving 400 Million Search Queries Monthly With Perplexity AI Perplexity AI, an AI-powered search engine, handles over 435 million monthly queries. Each query represents multiple AI inference requests. To meet this demand, the Perplexity AI team turned to NVIDIA H100 GPUs, Triton Inference Server and TensorRT-LLM.

Supporting over 20 AI models, including Llama 3 variations like 8B and 70B, Perplexity processes diverse tasks such as search, summarization and question-answering. By using smaller classifier models to route tasks to GPU pods, managed by NVIDIA Triton, the company delivers cost-efficient, responsive service under strict service level agreements.

Through model parallelism, which splits LLMs across GPUs, Perplexity achieved a threefold cost reduction while maintaining low latency and high accurac
LINK: https://blogs.nvidia.com/blog/ai-inference-platform/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

04/08/2026

Dalet Announces Commercial Availability of Dalia, Bringing Media-Aware Agentic AI to Enterprise Productions

Dalet, a leading technology and service provider for media-rich organizations, t...

04/07/2026

Detective Conan: Fallen Angel of the Highway Opens in Dolby Cinemas Across Japan, Presented in Dolby Atmos and Dolby ...

April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

07/05/2026

Colorfront Streaming Server Awarded Trusted Partner Network Gold Logo Certification

January 8, 2024 Colorfront (colorfront.com), a leader in high-performance, on-s...

07/05/2026

Colorfront Unleashes New Opportunities for Content Owners Plus Other Ground-Breaking Visual Experiences at Nab 2024

March 20, 2024 NAB 2024, Las Vegas - Colorfront (colorfront.com), the multi-awa...

07/05/2026

Colorfront Delivers Innovations for HDR Cinema at Cinemacon 2025

April 1, 2025 CINEMACON, APRIL 1, 2025 - Colorfront (colorfront.com), the multi-award-winning developer of high-performance dailies/transcoding/streaming syste...

07/05/2026

Colorfront to Showcase Advanced HDR Cinema Mastering Solutions at CineEurope 2025

June 15, 2025 Colorfront (colorfront.com), an Academy and Emmy Award-winning de...

07/05/2026

Colorfront at ICTA Barcelona Cinema Technology Summit

June 15, 2025 Colorfront participated in the ICTA Barcelona Cinema Technology Summit on Sunday, June 15, 2025. Held at The Phenomena Experience, the event feat...

07/05/2026

Colorfront Does IBC 2024 Live From Stockholm!

July 1, 2025 Colorfront (colorfront.com), the multi-award-winning developer of high-performance dailies/transcoding/streaming systems for motion pictures, OTT,...

07/05/2026

Colorfront Transkoder Delivers a Masterful Performance at Annapurna Studios

July 1, 2025 Passion and dedication will take you places. Come with us on a short trip to the heart of India, where Annapurna Studios is living-up to the inspi...

07/05/2026

Colorfront Opens a New Chapter in Color Tools For Cinema & Television

July 3, 2025 Colorfront (colorfront.com), the multi-award-winning developer of high-performance dailies/transcoding/streaming systems for motion pictures, OTT,...

07/05/2026

Colorfront at IBC 2025: Advanced Tools Make Easy Work of Critical Mastering Tasks and More!

September 1, 2025 IBC 2025, Amsterdam - Colorfront (colorfront.com) - the Acade...

07/05/2026

Colorfront Introduces Colorfront Immersive Utility, a New Mac App for Creating Apple Immersive Video

April 17, 2026 LOS ANGELES - April 17, 2026 - Colorfront today announced Colorf...

07/05/2026

Colorfront Delivers Even More AI Automation Power and Extends Technology Partnerships with Dolby Apple

April 23, 2026 NAB 2026, Las Vegas - the Academy and Emmy Award-winning develop...

07/05/2026

CNN Founder Ted Turner Dies at 87

Share Copy link Facebook X Linkedin Bluesky Email...

07/05/2026

FCC Urges Appeals Court to Toss Challenges to Nexstar-Tegna Deal

Share Copy link Facebook X Linkedin Bluesky Email...

07/05/2026

Carr Announces FCC Staff Promotions

Share Copy link Facebook X Linkedin Bluesky Email...

07/05/2026

Recreating the 1974 Doctor Who Time Tunnel in After Effects

Recreating the 1974 Doctor Who Time Tunnel in After Effects Graham Quince May 6, 2026 0 Comments The Time Tunnel from Doctor Who titles is one of th...

06/05/2026

Wisycom RF Solutions Support Gravity Medias Live Cycling and Marathon Broadcasts

Gravity Media Chief RF Communications Engineer Glenn Willems uses Wisycom RF over Fiber and wireless solutions across major cycling events and international mar...

06/05/2026

Sennheiser Spectera Module Now Available in Bitfocus Companion and Buttons

A Sennheiser Spectera module is now available in Bitfocus Companion and Buttons, enabling direct integration of Spectera with the two software platforms. The mo...

06/05/2026

Ted Turner, Cable Television Pioneer, Sports Broadcasting Hall of Famer, Dead at 87

Ted Turner, the visionary media entrepreneur whose appetite for disruption helpe...

06/05/2026

FIFA World Cup 2026: Peacock Launches Visin de Campo (aka Pitchside Live), Will Stream All 104 Matches in Spanish

Peacock is going all-in on the beautiful game - streaming all 104 FIFA World Cup...

06/05/2026

NoiseWorks Audio launch VoiceAssist Basic, Standard & Advanced

New pricing tiers for vocal/dialogue restoration tool NoiseWorks Audio's AI-powered vocal and dialogue processing plug-in is now available in three diff...

06/05/2026

RME TotalMix FX 2 now available

Popular mixing & routing software overhauled Following a recent public beta test, RME have launched the final release version of the powerful mixing and rou...

06/05/2026

Focusrite: Designing The ISA C8X Audio Interface

New SOS Video Feature Focusrites ISA C8X is a milestone product that brings together the companys analogue heritage and their expertise in digital audio. Yo...

06/05/2026

L3Harris to Boost Polish Navy Combat Power with Advanced Ship System

Polands Miecznik-class frigates are part of the largest contract in Polish shipbuilding history. (Image Credit: PGZ Stocznia Wojenna)...

06/05/2026

FCC's Anna Gomez Urges Rigorous Review of Paramount-WBD Merger

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Riedel Ups Marc Engroff to CFO, Shifts Frank Eischet to Group COO

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Amagi Launches In-Content Ads' to Attract More CTV Advertisers

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Riedel Expands Leadership Structure Appoints Marc Engroff...

Riedel Communications today announced the expansion of its leadership structure as part of a strategic initiative to strengthen both its operational management ...

06/05/2026

Production Sound Mixer Dirk Sciarrotta Delivers Camera Re...

For nearly three decades, Veteran Production Sound Mixer and Five-time Emmy Award Winner Dirk Sciarrotta has helped define the sonic identity of the long-runnin...

06/05/2026

ZEISS CinCraft LensCore: Cinema Lens Looks for Compositing

ZEISS CinCraft LensCore: Cinema Lens Looks for Compositing Brie Clayton May 6, 2026 0 Comments ZEISS announces the launch of CinCraft LensCore, a nove...

06/05/2026

Wisycom Solves Extreme RF Challenges Across Miles of Live Action for Gravity Media

Wisycom Solves Extreme RF Challenges Across Miles of Live Action for Gravity Med...

06/05/2026

NAB Launches Weekly Podcast on Local Broadcast Policy

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Mavis Launches Mavis Studio iPad For Media Production

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Narrative Entertainment partners with Encompass to provid...

Narrative Entertainment has partnered with Encompass to deliver high-quality subtitling of its Great! network content using the Altitude Intelligence AI assiste...

06/05/2026

SipRadius extends its seamless creation and connectivity...

SipRadius, widely recognized for making content processing and connectivity secure and seamless, is proud to launch a dramatic new approach to AI content creati...

06/05/2026

Big Blue Marble at ANGA COM - TV as a Service in the spot...

When the broadband and media industry gathers at ANGA COM in Cologne from May 19 to 21, Big Blue Marble will be at the forefront. The international broadcast an...

06/05/2026

Cinegy makes its MPTS debut with software-defined televis...

Cinegy GmbH, a leading developer of software-defined television technology, is proud to exhibit at MPTS for the first time. Visitors to the stand will discover ...

06/05/2026

Val Jeanty Receives 2026 Doris Duke Artist Award

Val Jeanty Receives 2026 Doris Duke Artist Award Jeanty, a composer, percussionist, and turntablist, is the fourth Berklee recipient of the prestigious award ...

06/05/2026

Zeiss Launches CinCraft LensCore

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Gomez Urges Rigorous FCC Review of Paramount-WBD Merger

Share Copy link Facebook X Linkedin Bluesky Email...

06/05/2026

Wisycom Solves Extreme RF Challenges Across Miles of Live...

When live cycling races and international marathons stretch for miles across cities and countryside, there is no margin for RF failure in live broadcast. As Chi...

06/05/2026

ZEISS CinCraft LensCore - Cinema Lens Looks for Compositi...

Oberkochen/Germany, May 5, 2026 ZEISS announces the launch of CinCraft LensCore, a novel solution for creating physically based cinematic lens looks for visual...

06/05/2026

VEON's Kyivstar Authorized to Resell Starlink for Businesses & Enterprises in Ukraine

06 May 2026 VEON's Kyivstar Authorized to Resell Starlink for Businesses & ...

06/05/2026

UKTV Brings the Early Years of Neighbours to U in 420Episode Back Catalogue Deal

UKTV has secured the exclusive rights to the early back catalogue of iconic Australian drama series Neighbours, following a landmark content deal with Fremantle...

06/05/2026

Sky commissions feature documentary to mark 10th anniversary of Manchester Arena bomb

Wednesday 6 May 2026 Sky commissions feature documentary to mark 10th anniversa...

06/05/2026

Sky and Formula 1 agree long-term partnership across UK, Ireland and Italy

Wednesday 6 May 2026 Sky and Formula 1 agree long-term partnership across UK, Ireland and Italy Sky in the UK & Ireland to remain the home of Formula 1 until ...