Sony Pixel Power calrec Sony

Mission NIMpossible: Decoding the Microservices That Accelerate Generative AI

10/07/2024

Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible and showcases new hardware, software, tools and accelerations for NVIDIA RTX PC and workstation users.

In the rapidly evolving world of artificial intelligence, generative AI is captivating imaginations and transforming industries. Behind the scenes, an unsung hero is making it all possible: microservices architecture.

The Building Blocks of Modern AI Applications Microservices have emerged as a powerful architecture, fundamentally changing how people design, build and deploy software.

A microservices architecture breaks down an application into a collection of loosely coupled, independently deployable services. Each service is responsible for a specific capability and communicates with other services through well-defined application programming interfaces, or APIs. This modular approach stands in stark contrast to traditional all-in-one architectures, in which all functionality is bundled into a single, tightly integrated application.

By decoupling services, teams can work on different components simultaneously, accelerating development processes and allowing updates to be rolled out independently without affecting the entire application. Developers can focus on building and improving specific services, leading to better code quality and faster problem resolution. Such specialization allows developers to become experts in their particular domain.

Services can be scaled independently based on demand, optimizing resource utilization and improving overall system performance. In addition, different services can use different technologies, allowing developers to choose the best tools for each specific task.

A Perfect Match: Microservices and Generative AI The microservices architecture is particularly well-suited for developing generative AI applications due to its scalability, enhanced modularity and flexibility.

AI models, especially large language models, require significant computational resources. Microservices allow for efficient scaling of these resource-intensive components without affecting the entire system.

Generative AI applications often involve multiple steps, such as data preprocessing, model inference and post-processing. Microservices enable each step to be developed, optimized and scaled independently. Plus, as AI models and techniques evolve rapidly, a microservices architecture allows for easier integration of new models as well as the replacement of existing ones without disrupting the entire application.

NVIDIA NIM: Simplifying Generative AI Deployment As the demand for AI-powered applications grows, developers face challenges in efficiently deploying and managing AI models.

NVIDIA NIM inference microservices provide models as optimized containers to deploy in the cloud, data centers, workstations, desktops and laptops. Each NIM container includes the pretrained AI models and all the necessary runtime components, making it simple to integrate AI capabilities into applications.

NIM offers a game-changing approach for application developers looking to incorporate AI functionality by providing simplified integration, production-readiness and flexibility. Developers can focus on building their applications without worrying about the complexities of data preparation, model training or customization, as NIM inference microservices are optimized for performance, come with runtime optimizations and support industry-standard APIs.

AI at Your Fingertips: NVIDIA NIM on Workstations and PCs Building enterprise generative AI applications comes with many challenges. While cloud-hosted model APIs can help developers get started, issues related to data privacy, security, model response latency, accuracy, API costs and scaling often hinder the path to production.

Workstations with NIM provide developers with secure access to a broad range of models and performance-optimized inference microservices.

By avoiding the latency, cost and compliance concerns associated with cloud-hosted APIs as well as the complexities of model deployment, developers can focus on application development. This accelerates the delivery of production-ready generative AI applications - enabling seamless, automatic scale out with performance optimization in data centers and the cloud.

The recently announced general availability of the Meta Llama 3 8B model as a NIM, which can run locally on RTX systems, brings state-of-the-art language model capabilities to individual developers, enabling local testing and experimentation without the need for cloud resources. With NIM running locally, developers can create sophisticated retrieval-augmented generation (RAG) projects right on their workstations.

Local RAG refers to implementing RAG systems entirely on local hardware, without relying on cloud-based services or external APIs.

Developers can use the Llama 3 8B NIM on workstations with one or more NVIDIA RTX 6000 Ada Generation GPUs or on NVIDIA RTX systems to build end-to-end RAG systems entirely on local hardware. This setup allows developers to tap the full power of Llama 3 8B, ensuring high performance and low latency.

By running the entire RAG pipeline locally, developers can maintain complete control over their data, ensuring privacy and security. This approach is particularly helpful for developers building applications that require real-time responses and high accuracy, such as customer-support chatbots, personalized content-generation tools and interactive virtual assistants.

Hybrid RAG combines local and cloud-based resources to optimize performance and flexibility in AI applications. With NVIDIA AI Workbench, developers can get started with the hybrid-RAG Workbench Project - an example application that can be used to run vector databases and embedding models locally whil
LINK: https://blogs.nvidia.com/blog/ai-decoded-nim/...
See more stories from nvidia

Most recent headlines

09/11/2025

Dalet Unveils Agentic AI Media Workflows at IBC2025

Dalet today announced a transformative leap forward for media operations: Agentic Artificial Intelligence (AI) that unifies the Dalet ecosystem under one natura...

18/10/2025

NESN Taps Harmonic for Primary Live Sports Distribution

New England Sports Network (NESN) has chosen Harmonic, working with Astound Business Solutions, as its enterprise technology partner to transform primary distri...

18/10/2025

DirecTV Launches Gray's Gulf Coast Sports & Entertainment Network

NEW ORLEANS, La. In the run-up to the start of the NBA season, WVUE-TV and Gray Local Media have announced a deal with DirecTV that will greatly expand access t...

18/10/2025

Berklee Celebrates 40 Years of the Fall Together Concert

Berklee Celebrates 40 Years of the Fall Together Concert Faculty composers Bob Pilkington and Greg Hopkins are among the featured artists for this year's ...

17/10/2025

NEP Group Receives New Equity Investment From 26North Partners LP, Co-Investors

NEP Group Receives New Equity Investment From 26North Partners LP, Co-InvestorsCarlyle remains the largest shareholder as the company prepares for the futureBy ...

17/10/2025

Apple Lands Five-Year Deal for F1 Distribution in the U.S.

Apple Lands Five-Year Deal for F1 Distribution in the U.S.Besides airing on Apple TV, the sport will be amplified on other Apple servicesBy Ken Kerschbaumer, Ed...

17/10/2025

SVG Sit-Down: Marshall Electronics' Bernie Keach on the Future of PTZ Cameras

SVG Sit-Down: Marshall Electronics' Bernie Keach on the Future of PTZ Camera...

17/10/2025

L2 Productions' REMI Facility in Austin Can Produce Content From Anywhere

L2 Productions' REMI Facility in Austin Can Produce Content From AnywhereMusic festivals, sports events are produced via flypacks and remote control roomsBy...

17/10/2025

Give Me the Backstory: Get to Know Sarah Dowland, the Filmmaker Behind Sue Bird: In The Clutch

By Lucy Spicer One of the most exciting things about the Sundance Film Festival...

17/10/2025

Cooper Raiff Returns to the Sundance Film Festival With His Independent Series Hal & Harper

(L-R) Christopher Meyer, Addison Timlin, Cooper Raiff, Lili Reinhart, Alyah Chan...

17/10/2025

Ferramenta de arte da capa de playlists do Spotify chega ao Brasil com uma noite de autoexpresso

M sica e arte se uniram em uma noite especial na semana passada na ZIV Gallery, ...

17/10/2025

Spotify's Custom Playlist Cover Art Tool Arrives in Brazil With a Night of Self-Expression

Music and art came together for one special night last week at ZIV Gallery, an i...

17/10/2025

Spotify and FC Barcelona Extend Partnership Through 2030

Spotify and FC Barcelona are extending our partnership through 2030, continuing a collaboration that's redefining how fans, players, and artists connect. Th...

17/10/2025

Sports Fishing Championship Deploys DigitalGlue Storage Platform

MURRIETA, Calif. The Sports Fishing Championship (SFC) has deployed DigitalGlue's creative.space storage platform to streamline video production by centrali...

17/10/2025

TV Ad Impressions for Football Spiked in Q3

BELLEVUE, Wash. Football continued to cement its reputation as a bulwark of TV advertising in Q3 2025 with new data from iSpot that showed both the NFL and coll...

17/10/2025

Reeling in the Chaos Sports Fishing Championship Simplifi...

The Sports Fishing Championship (SFC), the premier competitive saltwater fishing series, has transformed its production workflow by adopting creative.space, the...

17/10/2025

QuickLink Unveils StudioPro Version 4 With Major Enhancem...

QuickLink, a leading provider of award-winning multi-camera video productions and remote contribution solutions, announces the release of StudioPro Version 4, ...

17/10/2025

Westcoast Pixel dazzles with dynamic 3D video projections

Although the annual Grammy Awards celebration is best known for recognizing achievements in the recording industry, the show often proves a visual spectacle as ...

17/10/2025

Alex Dunfey Promoted to CTO at OpenDrives

OpenDrives, Inc., a leading provider of software-defined data storage and data services, has promoted Alex Dunfey to Chief Technology Officer (CTO) from his for...

17/10/2025

University of Arizona Scales Up Broadcast Capabilities Wi...

The University of Arizona (UofA) has significantly upgraded its broadcast communication infrastructure with the integration of Riedel Communications' advanc...

17/10/2025

NESN Redefines Regional Sports Video Delivery with Harmon...

Harmonic (NASDAQ: HLIT) today announced that New England Sports Network (NESN), owned by Fenway Sports Group and Delaware North, has selected Harmonic as its en...

17/10/2025

Austin PBS Expands Facility-Wide Production Communication...

Austin PBS has recently upgraded its facility-wide communications infrastructure, deploying Clear-Com 's Eclipse HX, FreeSpeak II beltpacks, and V-Series ...

17/10/2025

ZEISS Opens BETA Registration for CinCraft Virtual Lens T...

ZEISS announces an open call for the closed BETA testing phase of CinCraft Virtual Lens Technology, the innovative digital tool that brings authentic lens chara...

17/10/2025

Lightware powers hybrid learning transformation at Centri...

Situated in the town of Kokkola, Centria University of Applied Sciences offers higher education across five core fields: engineering, business, social and healt...

17/10/2025

Pebble to automate CobbTV

Public information channel in Georgia, USA, to implement a powerful, simple, and cost-effective playout automation platform. Pebble, the leading automation, co...

17/10/2025

HBO Maxs Global Expansion Surpasses 100 Market Milestone

HBO Max is reporting that it has launched in 15 new markets, including Bangladesh, Cambodia, Macau, Pakistan, Sri Lanka and Ukraine, boosting the streaming serv...

17/10/2025

Netflix Expands Into Video Podcasts With Spotify Deal

Netflix said it will make a major push into video podcasts, inking a wide-ranging deal with Spotify through which it will offer 16 podcasts in the U.S. starting...

17/10/2025

Viamedia Rebrands as Viamedia.ai

Lexington, Ky. As part of a push to highlight its advanced advertising capabilities, Viamedia has launched a new AI-powered ad tech platform and officially rebr...

17/10/2025

QuickLink to Showcase StudioPro Version 4 at NAB Show New York

NEW YORK QuickLink has announced the release of StudioPro Version 4, which the company is calling the most significant upgrade yet to its flagship video product...

17/10/2025

Apple, NBCU to Launch Apple TV, Peacock Streaming Bundles

NEW YORK and CUPERTINO, Calif. Apple and NBCUniversal said they will sell Apple TV and Peacock streaming bundles to U.S. subscribers starting Oct. 20....

17/10/2025

Q&A with Boston Conservatory Choral Conductor Stephen Spinelli

Q&A with Boston Conservatory Choral Conductor Stephen Spinelli How his research into the lost manuscripts of composer Florence Price led to a Grammy-winning c...

17/10/2025

Netflix ISP Speed Index for September 2025

Back to All News Netflix ISP Speed Index for September 2025 Product 17 October 2025 Global Link copied to clipboard This month, 1% of Internet Service Pro...

17/10/2025

Open Source AI Week - How Developers and Contributors Are Advancing AI Innovation

NVIDIA's on the ground at Open Source AI Week. Stay tuned for a celebration ...

17/10/2025

The Engines of American-Made Intelligence: NVIDIA and TSMC Celebrate First NVIDIA Blackwell Wafer Produced in the US

AI has ignited a new industrial revolution. NVIDIA and TSMC are working togethe...

17/10/2025

Showcasing global expertise in safety and risk management

Gexcon is a trusted safety and risk management partner for complex, high hazard environments. ICG has been a dedicated marketing partner to Gexcon since 2018, b...

17/10/2025

Kingfishr, David Walliams, Baz Ashmawy and Celine Byrne among the guests on this week's Late Late Show

Here is your host, Patrick Kielty! After an incredible breakthrough year, Kingf...

16/10/2025

SVG Sit-Down: FUJIFILM Execs on GFX ETERNA 55 Camera, Importance of Shallow-Depth-of-Field Production

SVG Sit-Down: FUJIFILM Execs on GFX ETERNA 55 Camera, Importance of Shallow-Dept...

16/10/2025

Squash's Most Ambitious Broadcast Production To Be Deployed at Comcast Business U.S. Open

Squash's Most Ambitious Broadcast Production To Be Deployed at Comcast Busin...

16/10/2025

Main Street Sports Group Inks Deal With Omaha Productions, Launches Original-Content Division

Main Street Sports Group Inks Deal With Omaha Productions, Launches Original-Con...

16/10/2025

A Historic Precursor? FIFA, HBS, DAZN Offer an Inside Look at Production of FIFA Club World Cup 2025

A Historic Precursor? FIFA, HBS, DAZN Offer an Inside Look at Production of FIFA...

16/10/2025

Prime Video Offers Sneak Peak at New NBA on Prime Studio

Prime Video Offers Sneak Peak at New NBA on Prime StudioThe massive 13,000-sq-ft, two-story studio features a LED regulation half court and hoopBy Jason Dachman...

16/10/2025

SVG Remote Production Forum Draws Record Crowd for Visit to PGA TOUR Studios, Deep Dive Into REMI Workflows

SVG Remote Production Forum Draws Record Crowd for Visit to PGA TOUR Studios, De...

16/10/2025

BitFire's Ben Grafchik on How Growing Cloud Workflows Are Impacting the Live-Sports World

BitFire's Ben Grafchik on How Growing Cloud Workflows Are Impacting the Live...

16/10/2025

CultureCon Uncut' Video Podcast Returns for Season 2 With Host Imani Ellis

In 2017, Imani Ellis launched CultureCon, a conference that's become a must-attend event for more than 10,000 diverse creatives and Black professionals to c...

16/10/2025

Fresh Spins on Holiday Standards: The 2025 Spotify Singles Have Arrived

It might still be a little early to break out the tinsel and mistletoe, but Spotify's already queuing up some holiday magic. This year's Spotify Singles...

16/10/2025

Spotify to Publish First Independent Author Releases Through Audiobook Selects, With More Planned Ahead

Earlier this year, our in-house publishing imprint, Spotify Audiobooks, put out ...

16/10/2025

L3Harris Integrates VAMPIRE Aboard GM Defense's Infantry Squad Vehicle

VAMPIRE has been integrated onto GM Defenses Infantry Squad Vehicle (ISV), providing a mobile solution to effectively and affordably counter small drone threat...

16/10/2025

Unseen and Unmatched: L3Harris Marks a First in Infrared Tracking Technology

The AgilePod mounted on the host aircraft....

16/10/2025

Consumer attitudes toward in-car entertainment and media preferences highlighted in new Gracenote automotive report

60% say infotainment systems are a critical purchasing or leasing consideration,...