SLMming Down Latency: How NVIDIA's First On-Device Small Language Model Makes Digital Humans More Lifelike
21/08/2024
At Gamescom this week, NVIDIA announced that NVIDIA ACE - a suite of technologies for bringing digital humans to life with generative AI - now includes the company's first on-device small language model (SLM), powered locally by RTX AI.
The model, called Nemotron-4 4B Instruct, provides better role-play, retrieval-augmented generation and function-calling capabilities, so game characters can more intuitively comprehend player instructions, respond to gamers, and perform more accurate and relevant actions.
Available as an NVIDIA NIM microservice for cloud and on-device deployment by game developers, the model is optimized for low memory usage, offering faster response times and providing developers a way to take advantage of over 100 million GeForce RTX-powered PCs and laptops and NVIDIA RTX-powered workstations.
The SLM Advantage An AI model's accuracy and performance depends on the size and quality of the dataset used for training. Large language models are trained on vast amounts of data, but are typically general-purpose and contain excess information for most uses.
SLMs, on the other hand, focus on specific use cases. So even with less data, they're capable of delivering more accurate responses, more quickly - critical elements for conversing naturally with digital humans.
Nemotron-4 4B was first distilled from the larger Nemotron-4 15B LLM. This process requires the smaller model, called a student, to mimic the outputs of the larger model, appropriately called a teacher. During this process, noncritical outputs of the student model are pruned or removed to reduce the parameter size of the model. Then, the SLM is quantized, which reduces the precision of the model's weights.
With fewer parameters and less precision, Nemotron-4 4B has a lower memory footprint and faster time to first token - how quickly a response begins - than the larger Nemotron-4 LLM while still maintaining a high level of accuracy due to distillation. Its smaller memory footprint also means games and apps that integrate the NIM microservice can run locally on more of the GeForce RTX AI PCs and laptops and NVIDIA RTX AI workstations that consumers own today.
This new, optimized SLM is also purpose-built with instruction tuning, a technique for fine-tuning models on instructional prompts to better perform specific tasks. This can be seen in Mecha BREAK, a video game in which players can converse with a mechanic game character and instruct it to switch and customize mechs.
ACEs Up ACE NIM microservices allow developers to deploy state-of-the-art generative AI models through the cloud or on RTX AI PCs and workstations to bring AI to their games and applications. With ACE NIM microservices, non-playable characters (NPCs) can dynamically interact and converse with players in the game in real time.
ACE consists of key AI models for speech-to-text, language, text-to-speech and facial animation. It's also modular, allowing developers to choose the NIM microservice needed for each element in their particular process.
NVIDIA Riva automatic speech recognition (ASR) processes a user's spoken language and uses AI to deliver a highly accurate transcription in real time. The technology builds fully customizable conversational AI pipelines using GPU-accelerated multilingual speech and translation microservices. Other supported ASRs include OpenAI's Whisper, a open-source neural net that approaches human-level robustness and accuracy on English speech recognition.
Once translated to digital text, the transcription goes into an LLM - such as Google's Gemma, Meta's Llama 3 or now NVIDIA Nemotron-4 4B - to start generating a response to the user's original voice input.
Next, another piece of Riva technology - text-to-speech - generates an audio response. ElevenLabs' proprietary AI speech and voice technology is also supported and has been demoed as part of ACE, as seen in the above demo.
Finally, NVIDIA Audio2Face (A2F) generates facial expressions that can be synced to dialogue in many languages. With the microservice, digital avatars can display dynamic, realistic emotions streamed live or baked in during post-processing.
The AI network automatically animates face, eyes, mouth, tongue and head motions to match the selected emotional range and level of intensity. And A2F can automatically infer emotion directly from an audio clip.
Finally, the full character or digital human is animated in a renderer, like Unreal Engine or the NVIDIA Omniverse platform.
AI That's NIMble In addition to its modular support for various NVIDIA-powered and third-party AI models, ACE allows developers to run inference for each model in the cloud or locally on RTX AI PCs and workstations.
The NVIDIA AI Inference Manager software development kit allows for hybrid inference based on various needs such as experience, workload and costs. It streamlines AI model deployment and integration for PC application developers by preconfiguring the PC with the necessary AI models, engines and dependencies. Apps and games can then orchestrate inference seamlessly across a PC or workstation to the cloud.
ACE NIM microservices run locally on RTX AI PCs and workstations, as well as in the cloud. Current microservices running locally include Audio2Face, in the Covert Protocol tech demo, and the new Nemotron-4 4B Instruct and Whisper ASR in Mecha BREAK.
To Infinity and Beyond Digital humans go far beyond NPCs in games. At last month's SIGGRAPH conference, NVIDIA previewed James, an interactive digital human that can connect with people using emotions, humor and more. James is based on
LINK: | https://blogs.nvidia.com/blog/ai-decoded-gamescom-ace-nemotron-instruc... |
See more stories from nvidia |
More from Nvidia
11/10/2024
NVIDIA AI Summit Panel Outlines Autonomous Driving Safety
The autonomous driving industry is shaped by rapid technological advancements and the need for standardization of guidelines to ensure the safety of both autono...
11/10/2024
Game-Changer: How the World's First GPU Leveled Up Gaming and Ignited the AI Era
In 1999, fans lined up at Blockbuster to rent chunky VHS tapes of The Matrix. Y2...
10/10/2024
The Next Chapter Awaits: Dive Into Diablo IV's' Latest Adventure Vessel of Hatred' on GeForce NOW
Prepare for a devilishly good time this GFN Thursday as the critically acclaimed...
10/10/2024
AI'll Be by Your Side: Mental Health Startup Enhances Therapist-Client Connections
Half of the world's population will experience a mental health disorder - bu...
09/10/2024
AI Summit: US Energy Secretary Highlights AI's Role in Science, Energy and Security
AI can help solve some of the world's biggest challenges - whether climate c...
09/10/2024
Flux and Furious: New Image Generation Model Runs Fastest on RTX AI PCs and Workstations
Editor's note: This post is part of the AI Decoded series, which demystifies...
09/10/2024
What's the ROI? Getting the Most Out of LLM Inference
Large language models and the applications they power enable unprecedented opportunities for organizations to get deeper insights from their data reservoirs and...
08/10/2024
NVIDIA AI Summit Highlights Game-Changing Energy Efficiency and AI-Driven Innovation
Accelerated computing is sustainable computing, Bob Pette, NVIDIA's vice pre...
08/10/2024
Accelerated Computing Key to Quantum Research
A recently released joint research paper by NVIDIA, Moderna and Yale reviews how techniques from quantum machine learning (QML) may enhance drug discovery metho...
08/10/2024
Pittsburgh Steels Itself for Innovation With Launch of NVIDIA AI Tech Community
Serving as a bridge for academia, industry and public-sector groups to partner on artificial intelligence innovation, NVIDIA is launching its inaugural AI Tech ...
08/10/2024
TSMC and NVIDIA Transform Semiconductor Manufacturing With Accelerated Computing
TSMC, the world leader in semiconductor manufacturing, is moving to production with NVIDIA's computational lithography platform, called cuLitho, to accelera...
08/10/2024
SETI Institute Researchers Engage in World's First Real-Time AI Search for Fast Radio Bursts
This summer, scientists supercharged their tools in the hunt for signs of life b...
08/10/2024
From Concept to Compliance, MITRE Digital Proving Ground Will Accelerate Validation of Autonomous Vehicles
The path to safe, widespread autonomous vehicles is going digital. MITRE - a go...
08/10/2024
A Not-So-Secret Agent: NVIDIA Unveils NIM Blueprint for Cybersecurity
Artificial intelligence is transforming cybersecurity with new generative AI tools and capabilities that were once the stuff of science fiction. And like many o...
08/10/2024
US Healthcare System Deploys AI Agents, From Research to Rounds
The U.S. healthcare system is adopting digital health agents to harness AI across the board, from research laboratories to clinical settings. The latest AI-acc...
07/10/2024
Foxconn to Build Taiwan's Fastest AI Supercomputer With NVIDIA Blackwell
NVIDIA and Foxconn are building Taiwan's largest supercomputer, marking a milestone in the island's AI advancement. The project, Hon Hai Kaohsiung Supe...
03/10/2024
No Tricks, Just Games: GeForce NOW Thrills With 22 Games in October
The air is crisp, the pumpkins are waiting to be carved, and GFN Thursday is ready to deliver some gaming thrills. GeForce NOW is unleashing a monster mash of ...
03/10/2024
How AI and Accelerated Computing Drive Energy Efficiency
AI isn't just about building smarter machines. It's about building a greener world. From optimizing energy use to reducing emissions, AI and accelerate...
02/10/2024
Brave New World: Leo AI and Ollama Bring RTX-Accelerated Local LLMs to Brave Browser Users
Editor's note: This post is part of the AI Decoded series, which demystifies...
01/10/2024
NVIDIA AI Summit DC: Industry Leaders Gather to Showcase AI's Real-World Impact
Washington, D.C., is where possibility has always met policy, and AI presents un...
27/09/2024
Bon Voyage: NIO Unveils ONVO L60 Smart Electric SUV, Built on NVIDIA DRIVE Orin
NIO's smart EV brand, ONVO, has unveiled the L60 flagship mid-size family SUV, built on the NVIDIA DRIVE Orin system-on-a-chip. Earlier this year, the auto...
26/09/2024
A Whole New World: GreedFall II: The Dying World' Joins GeForce NOW
Whether looking for a time-traveling adventure, strategic roleplay or epic action, anyone can find something to play on GeForce NOW, with over 2,000 games in th...
25/09/2024
Decoding How AI Can Accelerate Data Science Workflows
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
23/09/2024
To Save Lives, and Energy, Wellcome Sanger Institute Speeds Cancer Research With NVIDIA Accelerated Computing
The Wellcome Sanger Institute, a key contributor to the international Human Geno...
23/09/2024
NVIDIA Partners for Globally Inclusive AI in U.S. Government Initiative
NVIDIA is joining the U.S. government's launch of the Partnership for Global Inclusivity on AI (PGIAI), providing Deep Learning Institute training, GPU cred...
23/09/2024
High-Speed AI: Hitachi Rail Advances Real-Time Railway Analysis Using NVIDIA Technology
Hitachi Rail, a global transportation company powering railway systems in over 5...
20/09/2024
Medical Centers Tap AI, Federated Learning for Better Cancer Detection
A committee of experts from top U.S. medical centers and research institutes is harnessing NVIDIA-powered federated learning to evaluate the impact of federated...
19/09/2024
We've Fused Signal Processing and AI': NVIDIA CEO Outlines Future of Telecom at T-Mobile's Capital Markets Day
In a surprise appearance at T-Mobile's Capital Markets Day, NVIDIA founder a...
19/09/2024
Climate Week Forecast: Outlook Improving With AI, Accelerated Computing
All the electricity that powers NVIDIA's global operations will come from renewable sources by the end of January. It's the right fuel for the company&...
19/09/2024
FINAL FANTASY XVI' Soars Into the Cloud With GeForce NOW
GeForce NOW makes gamers' fantasies a reality by bringing top titles to the cloud. This week, the award-winning FINAL FANTASY XVI is available for members t...
18/09/2024
NVIDIA AI Aerial Launches to Optimize Wireless Networks, Deliver New Generative AI Experiences on One Platform
Telecommunications providers are transforming beyond voice and data services wit...
18/09/2024
How SonicJobs Uses AI Agents to Connect the Internet, Starting with Jobs
Companies in the US spend $15bn annually on talent acquisition. The most important metric in recruitment advertising is the conversion from the paid click on th...
17/09/2024
New AI Innovation Hub in Tunisia Drives Technological Advancement Across Africa
A new AI innovation hub for developers across Tunisia launched today in Novation City, a technology park that's designed to cultivate a vibrant, innovation ...
17/09/2024
Upgrade Livestreams With Twitch Enhanced Broadcasting and the NVIDIA Encoder
At TwitchCon - a global convention for the Twitch livestreaming platform-livestreamers and content creators this week can experience the latest technologies for...
12/09/2024
GeForce NOW to Bring Dead Rising Deluxe Remaster' to the Cloud at Launch
Rise and shine - Capcom's latest action-adventure game, Dead Rising Deluxe Remaster, heads to the cloud at launch next week. It's part of nine new titl...
11/09/2024
AI on the Air: Behind the Scenes at IBC With Holoscan for Media
AI is transforming the broadcast industry by enhancing the way content is created, distributed and consumed - but integrating the technology can be challenging....
11/09/2024
NVIDIA and Oracle to Accelerate AI and Data Processing for Enterprises
Enterprises are looking for increasingly powerful compute to support their AI workloads and accelerate data processing. The efficiency gained can translate to b...
11/09/2024
Ready to Roll: Nuro to License Its Autonomous Driving System
To accelerate autonomous vehicle development and deployment timelines, Nuro announced today it will license its Nuro Driver autonomous driving system directly t...
09/09/2024
Live Media Reimagined: NVIDIA Holoscan for Media Now Available for Production
Companies in broadcast, sports and streaming are transitioning to software-defined infrastructure to benefit from flexible deployment and to more easily adopt t...
06/09/2024
How AI Is Personalizing Customer Service Experiences Across Industries
Customer service departments across industries are facing increased call volumes, high customer service agent turnover, talent shortages and shifting customer e...
05/09/2024
19 New Games to Drop for GeForce NOW in September
Fall will be here soon, so leaf it to GeForce NOW to bring the games, with 19 joining the cloud in September. Get started with the seven games available to str...
05/09/2024
Three Ways to Ride the Flywheel of Cybersecurity AI
The business transformations that generative AI brings come with risks that AI itself can help secure in a kind of flywheel of progress. Companies who were qui...
04/09/2024
Volvo Cars EX90 SUV Rolls Out, Built on NVIDIA Accelerated Computing and AI
Volvo Cars' new, fully electric EX90 is making its way from the automaker's assembly line in Charleston, South Carolina, to dealerships around the U.S. ...
04/09/2024
Do the Math: New RTX AI PC Hardware Delivers More AI, Faster
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...
04/09/2024
Hammer Time: Machina Labs' Edward Mehr on Autonomous Blacksmith Bots and More
Edward Mehr works where AI meets the anvil. The company he cofounded, Machina L...
04/09/2024
Manufacturing Intelligence: Deltia AI Delivers Assembly Line Gains With NVIDIA Metropolis and Jetson
It all started at Berlin's Merantix venture studio in 2022, when Silviu Homo...
29/08/2024
From RAG to Richness: Startup Uplevels Retrieval-Augmented Generation for Enterprises
Well before OpenAI upended the technology industry with its release of ChatGPT i...
29/08/2024
Crystal-Clear Gaming: Visions of Mana' Sharpens on GeForce NOW
It's time to mana-fest the spirit of adventure with Square Enix's highly anticipated action role-playing game, Visions of Mana, launching today in the c...
28/08/2024
NVIDIA Blackwell Sets New Standard for Generative AI in MLPerf Inference Debut
As enterprises race to adopt generative AI and bring new services to market, the demands on data center infrastructure have never been greater. Training large l...
28/08/2024
More Than Fine: Multi-LoRA Support Now Available in NVIDIA RTX AI Toolkit
Editor's note: This post is part of the AI Decoded series, which demystifies AI by making the technology more accessible, and showcases new hardware, softwa...