Sony Pixel Power calrec Sony

Seamless in Seattle: NVIDIA Research Showcases Advancements in Visual Generative AI at CVPR

17/06/2024

NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.

More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.

NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.

NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.

Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.

Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.

At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.

Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.

Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.

JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.

JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.

https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4

New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.

The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.

FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.

NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.

Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.

VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.

The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.

VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
LINK: https://blogs.nvidia.com/blog/visual-generative-ai-cvpr-research/...
See more stories from nvidia

North America Stories

07/04/2026

Avid to Debut Avid Content Core on AWS at 2026 NAB Show

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

Frequency Launches Smarter Ways to Operate Streaming Chan...

Frequency, the engine behind many of the world's leading streaming television channels, at NAB 2026 will be launching new Studio services to help content ow...

07/04/2026

Ikegami to Introduce Expanded Range of Broadcast Producti...

Ikegami USA has chosen NAB 2026 in Las Vegas as the launch platform for new additions to its range of broadcast-quality television production equipment. These w...

07/04/2026

Kiloview Advancing Broadcast IP Workflows with a Smarter...

April 19, 2026, Las Vegas Kiloview, an innovative provider of AV-over-IP technologies, will showcase its latest broadcast IP solutions at NAB 2026, presenting...

07/04/2026

Bitmovin Adds Support for SGAI in its Playback Products t...

Bitmovin has announced support for Server-Guided Ad Insertion (SGAI) across its playback products using HLS interstitials, enabling more advanced ad-supported s...

07/04/2026

Synamedia unveils AI by Quortex - a just-in-time AI-plugi...

Synamedia is unveiling AI by Quortex at The NAB show, a just-in-time AI plugin framework that applies intelligence only when needed across video processing, dis...

07/04/2026

Cuez Brings Four New Innovations to NAB 2026 From Story-C...

Cuez will showcase four additions to its cloud-based newsroom, rundown and automation platform at NAB Show 2026 (April 18 22, Las Vegas, Booth N1867): Cuez ...

07/04/2026

Barix Extends Transport Options for Multi-Engine IP Encoder

Barix Extends Transport Options for Multi-Engine IP Encoder Brie Clayton April 7, 2026 0 Comments New for NAB, Barix adds SRT and RIST support to Mult...

07/04/2026

Elite Media Technologies Selects Interra Systems' BATON File-Based QC Solution

Elite Media Technologies Selects Interra Systems' BATON File-Based QC Soluti...

07/04/2026

Tightrope Media Systems to Debut Cablecast LiveBridge for Simultaneous Streaming at NAB 2026

Tightrope Media Systems to Debut Cablecast LiveBridge for Simultaneous Streaming...

07/04/2026

Cuez Brings Four New Innovations to NAB 2026: From Story-Centric Newsroom to Open AI Agent Framework

Cuez Brings Four New Innovations to NAB 2026: From Story-Centric Newsroom to Ope...

07/04/2026

ASG Names Andrea Cummis VP of Systems Engineering

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

KTVJ Completes Major Signal Upgrade

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

Hearst's WDSU to Air Million Dollar Rodeo Competition

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

Grass Valley Launches Future Playmakers Program

Share Copy link Facebook X Linkedin Bluesky Email...

07/04/2026

Saranyu Technologies Launches MATCH - Multi-View Sports S...

Designed for synchronized multi-stream playback, low-latency delivery, and real-time analytics, MATCH introduces a unified viewing experience for sports broadca...

07/04/2026

Berklee Students to Honor George Martin with Performance of Original Scores

Berklee Students to Honor George Martin with Performance of Original Scores The orchestra, led by associate professor Xander Rovang, will perform several work...

06/04/2026

Fab Five' Reunion Drives TNT and CBS's Experimental Final Four Altcast Built on REMI Workflow

Michigan legends bring a new voice to the broadcast as TNT Sports and CBS Sports...

06/04/2026

SVG New Sponsor Spotlight: Optikka CEO Daniel Evans on Scaling Sports Content with Programmatic Graphics

From high school sports all the way up to the major leagues, building high-quali...

06/04/2026

Quickplay and TwelveLabs Join AWS Business Outcomes Xcelerator Program

Quickplay, an AI company for the media and entertainment industry, has been accepted into the Advanced tier of the TwelveLabs Ecosystem Partner Program. Quickpl...

06/04/2026

Grass Valley Launches Future Playmakers Program for Students in Sports Production and Media Technology

Grass Valley has announced the Future Playmakers Program, a global initiative to...

06/04/2026

SVG All-Stars: Raasean Robinson, Gerente de Posproduccin y Operaciones de Estudio, FOX Deportes

El l der de operaciones impulsa la producci n en estudio mientras encuentra insp...

06/04/2026

SVG All-Stars: Raasean Robinson, Manager, Post Production and Studio Operations, FOX Deportes

The ops leader helps lead the charge in studio for the Spanish-language broadcas...

06/04/2026

Behind The Mic: SiriusXM Shares 2026 Masters Broadcast Team; ESPN to Produce Over 140+ Hours of Masters Live Coverage

Behind The Mic provides a roundup of recent news regarding on-air talent, includ...

06/04/2026

NHL Opens Innovation Lab in Partnership with Verizon, New Jersey Devils

The National Hockey League (NHL), in partnership with Verizon and the New Jersey Devils, today announced the opening of the NHL Innovation Lab powered by Verizo...

06/04/2026

ESPN+ To Stream Inaugural Rock League Curling Season

Rock League, a new professional curling league, has announced that ESPN+ will stream its inaugural 2026 season for fans in the United States. The first Rock Lea...

06/04/2026

ASG Appoints Andrea Cummis as VP of Systems Design and Engineering

Advanced Systems Group has announced the appointment of Andrea (Andy) Cummis as Vice President of Systems Design and Engineering. In this role, she will lead de...

06/04/2026

Source Media Group Launches Source Golf, a Creator-Driven YouTube Network Targeting Next-Gen Fans

Backed by Bolt Ventures, the venture brings Bryson DeChambeau, Grant Horvat, and...

06/04/2026

How the NHL's Innovation Lab Will Take Broadcast, Fan, and Team Tech to New Heights

With this environment we can start that collaboration even earlier because we ca...

06/04/2026

Baseball 2026: More AI, Better Viewing Choices

Share Copy link Facebook X Linkedin Bluesky Email...

06/04/2026

JB&A Announces Details for its Pre-NAB 2026 Event

Share Copy link Facebook X Linkedin Bluesky Email...

06/04/2026

Dalet Showcases Dalia Agentic AI and End-to-End Media Workflows at NAB Show 2026

Dalet Showcases Dalia Agentic AI and End-to-End Media Workflows at NAB Show 2026 Brie Clayton April 6, 2026 0 Comments Dalet, a leading technology and...

06/04/2026

OpenDrives Shows Off Sports Expertise in Sports Business Hub located in NAB Show's West Hall

OpenDrives Shows Off Sports Expertise in Sports Business Hub located in NAB Show...

06/04/2026

Proton to Demonstrate 3D Application at NAB 2026

Proton to Demonstrate 3D Application at NAB 2026 Brie Clayton April 6, 2026 0 Comments Yet further creative potential unleashed through innovation in ...

06/04/2026

Autoscript Highlights Voice-Driven Prompting and PTZ Solutions at NAB 2026

Autoscript Highlights Voice-Driven Prompting and PTZ Solutions at NAB 2026 Brie Clayton April 6, 2026 0 Comments Experience Autoscript Voice, PTZ prom...

06/04/2026

Mediaproxy Highlights Significant Enhancements to its LogServer suite at NAB Show 2026

Mediaproxy Highlights Significant Enhancements to its LogServer suite at NAB Sho...

06/04/2026

Tribeca Studios And Lilly Announce Winners Of Inaugural Vital Stories Filmmaker Program

April 6th, 2026 TRIBECA STUDIOS AND LILLY ANNOUNCE WINNERS OF INAUGURAL VITAL...

06/04/2026

Netflix Expands Kids Entertainment Lineup With Playground App for Games, New Shows & Returning Favorites

Back to All News Netflix Expands Kids Entertainment Lineup With Playground App ...

04/04/2026

Don't Be Lame: Arizona Men's Basketball Social Team Aims To Catch the Attention of Wildcats Fans

The University of Arizona's Men's Basketball team has only loss twice th...

04/04/2026

HDR Makes Its Men's Final Four Debut as CBS Sports and TNT Sports Collaborate on New Camera Tools and an IP-Powered Compound

1080p HDR arrives, a new generation of storytelling tools takes center stage, an...

04/04/2026

Fab Five Reunion Drives TNT and CBS's Experimental Final Four Altcast Built on REMI Workflow

Michigan legends bring a new voice to the broadcast as TNT Sports and CBS Sports...

04/04/2026

Sinclair to FCC: Broadcast Sports Drives Investment in Local News

Share Copy link Facebook X Linkedin Bluesky Email...

04/04/2026

Study: Worldwide Telecom Capex to Decline in 2026,

Share Copy link Facebook X Linkedin Bluesky Email...

04/04/2026

Ateme Delivers Full End-to-End Streaming Platform to Moldtelecom

Share Copy link Facebook X Linkedin Bluesky Email...

04/04/2026

FCC Plans Spending, Regulatory Fee Revenue Reductions in FY 2027

Share Copy link Facebook X Linkedin Bluesky Email...

04/04/2026

DHD Introduces AI-Based Audio Noise Reduction to XD3 IP Core

DHD Introduces AI-Based Audio Noise Reduction to XD3 IP Core Brie Clayton April 3, 2026 0 Comments The accompanying image shows the rear panel of the ...

04/04/2026

Macnica Redefines ST 2110 Flexibility with Two Speeds on One Card

Macnica Redefines ST 2110 Flexibility with Two Speeds on One Card Brie Clayton April 3, 2026 0 Comments New for NAB Show 2026, MEP100 SmartNIC now sup...

04/04/2026

Unified Media Workflows for Story-Centric Production

Unified Media Workflows for Story-Centric Production Brie Clayton April 3, 2026 0 Comments Framelight X unifies field capture, editing and publishing ...