Sony Pixel Power calrec Sony

Mission NIMpossible: Decoding the Microservices That Accelerate Generative AI

10/07/2024

Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible and showcases new hardware, software, tools and accelerations for NVIDIA RTX PC and workstation users.

In the rapidly evolving world of artificial intelligence, generative AI is captivating imaginations and transforming industries. Behind the scenes, an unsung hero is making it all possible: microservices architecture.

The Building Blocks of Modern AI Applications Microservices have emerged as a powerful architecture, fundamentally changing how people design, build and deploy software.

A microservices architecture breaks down an application into a collection of loosely coupled, independently deployable services. Each service is responsible for a specific capability and communicates with other services through well-defined application programming interfaces, or APIs. This modular approach stands in stark contrast to traditional all-in-one architectures, in which all functionality is bundled into a single, tightly integrated application.

By decoupling services, teams can work on different components simultaneously, accelerating development processes and allowing updates to be rolled out independently without affecting the entire application. Developers can focus on building and improving specific services, leading to better code quality and faster problem resolution. Such specialization allows developers to become experts in their particular domain.

Services can be scaled independently based on demand, optimizing resource utilization and improving overall system performance. In addition, different services can use different technologies, allowing developers to choose the best tools for each specific task.

A Perfect Match: Microservices and Generative AI The microservices architecture is particularly well-suited for developing generative AI applications due to its scalability, enhanced modularity and flexibility.

AI models, especially large language models, require significant computational resources. Microservices allow for efficient scaling of these resource-intensive components without affecting the entire system.

Generative AI applications often involve multiple steps, such as data preprocessing, model inference and post-processing. Microservices enable each step to be developed, optimized and scaled independently. Plus, as AI models and techniques evolve rapidly, a microservices architecture allows for easier integration of new models as well as the replacement of existing ones without disrupting the entire application.

NVIDIA NIM: Simplifying Generative AI Deployment As the demand for AI-powered applications grows, developers face challenges in efficiently deploying and managing AI models.

NVIDIA NIM inference microservices provide models as optimized containers to deploy in the cloud, data centers, workstations, desktops and laptops. Each NIM container includes the pretrained AI models and all the necessary runtime components, making it simple to integrate AI capabilities into applications.

NIM offers a game-changing approach for application developers looking to incorporate AI functionality by providing simplified integration, production-readiness and flexibility. Developers can focus on building their applications without worrying about the complexities of data preparation, model training or customization, as NIM inference microservices are optimized for performance, come with runtime optimizations and support industry-standard APIs.

AI at Your Fingertips: NVIDIA NIM on Workstations and PCs Building enterprise generative AI applications comes with many challenges. While cloud-hosted model APIs can help developers get started, issues related to data privacy, security, model response latency, accuracy, API costs and scaling often hinder the path to production.

Workstations with NIM provide developers with secure access to a broad range of models and performance-optimized inference microservices.

By avoiding the latency, cost and compliance concerns associated with cloud-hosted APIs as well as the complexities of model deployment, developers can focus on application development. This accelerates the delivery of production-ready generative AI applications - enabling seamless, automatic scale out with performance optimization in data centers and the cloud.

The recently announced general availability of the Meta Llama 3 8B model as a NIM, which can run locally on RTX systems, brings state-of-the-art language model capabilities to individual developers, enabling local testing and experimentation without the need for cloud resources. With NIM running locally, developers can create sophisticated retrieval-augmented generation (RAG) projects right on their workstations.

Local RAG refers to implementing RAG systems entirely on local hardware, without relying on cloud-based services or external APIs.

Developers can use the Llama 3 8B NIM on workstations with one or more NVIDIA RTX 6000 Ada Generation GPUs or on NVIDIA RTX systems to build end-to-end RAG systems entirely on local hardware. This setup allows developers to tap the full power of Llama 3 8B, ensuring high performance and low latency.

By running the entire RAG pipeline locally, developers can maintain complete control over their data, ensuring privacy and security. This approach is particularly helpful for developers building applications that require real-time responses and high accuracy, such as customer-support chatbots, personalized content-generation tools and interactive virtual assistants.

Hybrid RAG combines local and cloud-based resources to optimize performance and flexibility in AI applications. With NVIDIA AI Workbench, developers can get started with the hybrid-RAG Workbench Project - an example application that can be used to run vector databases and embedding models locally whil
LINK: https://blogs.nvidia.com/blog/ai-decoded-nim/...
See more stories from nvidia

More from Nvidia

26/06/2025

Run Google DeepMind's Gemma 3n on NVIDIA Jetson and RTX

As of today, NVIDIA now supports the general availability of Gemma 3n on NVIDIA RTX and Jetson. Gemma, previewed by Google DeepMind at Google I/O last month, in...

26/06/2025

Into the Omniverse: World Foundation Models Advance Autonomous Vehicle Simulation and Safety

Editor's note: This blog is a part of Into the Omniverse, a series focused o...

26/06/2025

Startup Uses NVIDIA RTX-Powered Generative AI to Make Coolers, Cooler

Mark Theriault founded the startup FITY envisioning a line of clever cooling products: cold drink holders that come with freezable pucks to keep beverages cold ...

26/06/2025

Game On With GeForce NOW, the Membership That Keeps on Delivering

This GFN Thursday rolls out a new reward and games for GeForce NOW members. Whether hunting for hot new releases or rediscovering timeless classics, members can...

24/06/2025

Introducing NVFP4 for Efficient and Accurate Low-Precision Inference

To get the most out of AI, optimizations are critical. When developers think about optimizing AI models for inference, model compression techniques-such as quan...

24/06/2025

HPE and NVIDIA Debut AI Factory Stack to Power Next Industrial Shift

To speed up AI adoption across industries, HPE and NVIDIA today launched new AI factory offerings at HPE Discover in Las Vegas. The new lineup includes everyth...

24/06/2025

NVIDIA and Partners Highlight Next-Generation Robotics, Automation and AI Technologies at Automatica

From the heart of Germany's automotive sector to manufacturing hubs across F...

19/06/2025

Step Inside the Vault: The Borderland' Series Arrives on GeForce NOW

GeForce NOW is throwing open the vault doors to welcome the legendary Borderland series to the cloud. Whether a seasoned Vault Hunter or new to the mayhem of P...

18/06/2025

Plug and Play: Build a G-Assist Plug-In Today

Project G-Assist - available through the NVIDIA App - is an experimental AI assistant that helps tune, control and optimize NVIDIA GeForce RTX systems. NVIDIA&...

17/06/2025

Hexagon Taps NVIDIA Robotics and AI Software to Build and Deploy AEON, a New Humanoid

As a global labor shortage leaves 50 million positions unfilled across industrie...

13/06/2025

NVIDIA and Deutsche Telekom Partner to Advance Germany's Sovereign AI

Industrial AI isn't slowing down. Germany is ready. Following London Tech Week and GTC Paris at VivaTech, NVIDIA founder and CEO Jensen Huang's Europea...

12/06/2025

NVIDIA TensorRT Boosts Stable Diffusion 3.5 Performance on NVIDIA GeForce RTX and RTX PRO GPUs

Generative AI has reshaped how people create, imagine and interact with digital ...

12/06/2025

Turn RTX ON With 40% Off Performance Day Passes

Level up GeForce NOW experiences this summer with 40% off Performance Day Passes. Enjoy 24 hours of premium cloud gaming with RTX ON, delivering low latency and...

11/06/2025

NVIDIA DRIVE Full-Stack Autonomous Vehicle Software Rolls Out

NVIDIA is launching a comprehensive, industry-defining autonomous vehicle (AV) software platform to accelerate large-scale deployment of safe, intelligent trans...

11/06/2025

NVIDIA Research Casts New Light on Scenes With AI-Powered Rendering for Physical AI Development

NVIDIA Research has developed an AI light switch for videos that can turn daytim...

11/06/2025

European Researchers Develop AI-Native Wireless Networks With NVIDIA 6G Research Portfolio

Using NVIDIA platforms, tools and libraries, European telecommunications institu...

11/06/2025

NVIDIA Scores Consecutive Win for End-to-End Autonomous Driving Grand Challenge at CVPR

NVIDIA was today named an Autonomous Grand Challenge winner at the Computer Visi...

11/06/2025

European Robot Makers Adopt NVIDIA Isaac, Omniverse and Halos to Develop Safe, Physical AI-Driven Robot Fleets

In the face of growing labor shortages and need for sustainability, European man...

11/06/2025

Retail Reboot: Major Global Brands Transform End-to-End Operations With NVIDIA

AI is packing and shipping efficiency for the retail and consumer packaged goods (CPG) industries, with a majority of surveyed companies in the space reporting ...

11/06/2025

NVIDIA Brings Physical AI to European Cities With New Blueprint for Smart City AI

Urban populations are expected to double by 2050, which means around 2.5 billion...

11/06/2025

Calling on LLMs: New NVIDIA AI Blueprint Helps Automate Telco Network Configuration

Telecom companies last year spent nearly $295 billion in capital expenditures an...

11/06/2025

European Broadcasting Union and NVIDIA Partner on Sovereign AI to Support Public Broadcasters

In a new effort to advance sovereign AI for European public service media, NVIDI...

11/06/2025

NVIDIA CEO Drops the Blueprint for Europe's AI Boom

At GTC Paris - held alongside VivaTech, Europe's largest tech event - NVIDIA founder and CEO Jensen Huang delivered a clear message: Europe isn't just a...

10/06/2025

The Blue Lion Supercomputer Will Run on NVIDIA Vera Rubin - Here's Why That Matters

Germany's Leibniz Supercomputing Centre, LRZ, is gaining a new supercomputer...

10/06/2025

Clear Skies Ahead: New NVIDIA Earth-2 Generative AI Foundation Model Simulates Global Climate at Kilometer-Scale Resolution

With a more detailed simulation of the Earth's climate, scientists and resea...

10/06/2025

Cisco and NVIDIA Advance Security for Enterprise AI Factories

Cisco and NVIDIA are helping set a new standard for secure, scalable and high-performance enterprise AI. Announced today at the Cisco Live conference in San Di...

09/06/2025

UK Prime Minister, NVIDIA CEO Set the Stage as AI Lights Up Europe

AI isn't waiting. And this week, neither is Europe. At London's Olympia, under a ceiling of steel beams and enveloped by the thrum of startup pitches, ...

08/06/2025

AI Maker, Not an AI Taker': UK Builds Its Vision With NVIDIA Infrastructure

U.K. Prime Minister Keir Starmer's ambition for Britain to be an AI maker, not an AI taker, is becoming a reality at London Tech Week. With NVIDIA's ...

05/06/2025

GeForce NOW Kicks Off a Summer of Gaming With 25 New Titles This June

GeForce NOW is a gamer's ticket to an unforgettable summer of gaming. With 25 titles coming this month and endless ways to play, the summer is going to be e...

04/06/2025

NVIDIA Blackwell Delivers Breakthrough Performance in Latest MLPerf Training Results

NVIDIA is working with companies worldwide to build out AI factories - speeding ...

04/06/2025

How 1X Technologies' Robots Are Learning to Lend a Helping Hand

Humans learn the norms, values and behaviors of society from each other - and Bernt B rnich, founder and CEO of 1X Technologies, thinks robots should learn like...

04/06/2025

NVIDIA RTX Blackwell GPUs Accelerate Professional-Grade Video Editing

4:2:2 cameras - capable of capturing double the color information compared with most standard cameras - are becoming widely available for consumers. At the same...

02/06/2025

Bring Receipts: New NVIDIA AI Blueprint Detects Fraudulent Credit Card Transactions With Precision

Editor's note: This blog, originally published on October 28, 2024, has been...

02/06/2025

Researchers and Students in Trkiye Build AI, Robotics Tools to Boost Disaster Readiness

Since a 7.8-magnitude earthquake hit Syria and T rkiye two years ago - leaving 5...

29/05/2025

The Supercomputer Designed to Accelerate Nobel-Worthy Science

Ready for a front-row seat to the next scientific revolution? That's the idea behind Doudna - a groundbreaking supercomputer announced today at Lawrence Be...

29/05/2025

Run LLMs on AnythingLLM Faster With NVIDIA RTX AI PCs

Large language models (LLMs), trained on datasets with billions of tokens, can generate high-quality content. They're the backbone for many of the most popu...

29/05/2025

RTX on Deck: The GeForce NOW Native App for Steam Deck Is Here

GeForce NOW is supercharging Valve's Steam Deck with a new native app - delivering the high-quality GeForce RTX-powered gameplay members are used to on a po...

28/05/2025

NVIDIA's Bartley Richardson on How Teams of AI Agents Provide Next-Level Automation

Building effective agentic AI systems requires rethinking how technology interac...

27/05/2025

How Dell Technologies Is Building the Engines of AI Factories With NVIDIA Blackwell

Over a century ago, Henry Ford pioneered the mass production of cars and engines...

27/05/2025

NVIDIA and Google Partnership Gains Momentum With the Latest Blackwell and Gemini Announcements

NVIDIA and Google share a long-standing relationship rooted in advancing AI inno...

22/05/2025

Sale Into Summer With 40% Off GeForce NOW Six-Month Performance Memberships

GeForce NOW is turning up the heat this summer with a hot new deal. For a limited time, save 40% on six-month Performance memberships and enjoy premium GeForce ...

21/05/2025

NVIDIA and SAP Bring AI Agents to the Physical World

As robots increasingly make their way to the largest enterprises' manufacturing plants and warehouses, the need for access to critical business and operatio...

20/05/2025

Siemens Makes Factory Floors Smarter With Industrial AI

Industrial AI is transforming how factories operate, innovate and scale. The convergence of AI, simulation and digital twins is poised to unlock new levels of ...

19/05/2025

NVIDIA and Microsoft Accelerate Agentic AI Innovation, From Cloud to PC

Agentic AI is redefining scientific discovery and unlocking research breakthroughs and innovations across industries. Through deepened collaboration, NVIDIA and...

19/05/2025

NVIDIA Research Breakthroughs Put Advanced Robots in Motion

Across robot training and development, NVIDIA Research is uncovering breakthroughs in areas such as multimodal generative AI and synthetic data generation. The...

19/05/2025

NVIDIA and Microsoft Advance Development on RTX AI PCs

Generative AI is transforming PC software into breakthrough experiences - from digital humans to writing assistants, intelligent agents and creative tools. NVI...

18/05/2025

NVIDIA CEO Envisions AI Infrastructure Industry Worth Trillions of Dollars'

Electricity. The Internet. Now it's time for another major technology, AI, to sweep the globe. NVIDIA founder and CEO Jensen Huang took the stage at a pack...

18/05/2025

NVIDIA Expands Omniverse Blueprint for AI Factory Digital Twins With New Ecosystem Integrations, Development Tools

Empowering engineering teams with more tools for building AI factories, NVIDIA t...

18/05/2025

AI Blueprint for Video Search and Summarization Now Available to Deploy Video Analytics AI Agents Across Industries

The age of video analytics AI agents is here. Video is one of the defining feat...

18/05/2025

Semiconductor Industry Accelerates Design Manufacturing With NVIDIA Blackwell and CUDA-X

TSMC, Cadence, KLA, Siemens and Synopsys are advancing semiconductor manufacturi...