Sony Pixel Power calrec Sony

Seamless in Seattle: NVIDIA Research Showcases Advancements in Visual Generative AI at CVPR

17/06/2024

NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.

More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.

NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.

NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.

Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.

Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.

At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.

Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.

Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.

JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.

JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.

https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4

New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.

The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.

FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.

NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.

Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.

VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.

The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.

VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
LINK: https://blogs.nvidia.com/blog/visual-generative-ai-cvpr-research/...
See more stories from nvidia

North America Stories

23/12/2025

How guilas Cibaeas Dominican Winter League Games Are Locally Produced for Global Audience

How guilas Cibae as Dominican Winter League Games Are Locally Produced for Glob...

23/12/2025

CAMB.AI Enables European Athletics to Offer Multi-Language Support

CAMB.AI Enables European Athletics to Offer Multi-Language SupportPlan is to eventually offer translation into all languages spoken in EuropeBy Ken Kerschbaumer...

23/12/2025

Analysis: As Sports Media Values Trend Negative, Scarcity and Quality Are King

Analysis: As sports media values trend negative, scarcity and quality are king By Callum McCarthy, Editor-at-Large Monday, December 22, 2025 - 14:08 Print ...

23/12/2025

ESPN, Disney, and NBA Return to the Animated Altcast Fray With Second Edition of Dunk the Halls'

ESPN, Disney, and NBA Return to the Animated Altcast Fray With Second Edition of...

23/12/2025

End the Year on a High Note and Donate to the Sports Broadcasting Fund Today!

End the Year on a High Note and Donate to the Sports Broadcasting Fund Today! By Ken Kerschbaumer, Editorial Director Tuesday, December 23, 2025 - 12:25 pm ...

23/12/2025

L3Harris Receives Letter of Intent from Kratos Defense for Production of Large Hypersonic Solid Rocket Motors

A Zeus motor is hot fire tested at L3Harris' Camden, Arkansas, solid rocket ...

23/12/2025

FCC Bans All New Foreign-Made Drones

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Gray Media Renews Its NBC Affiliation Agreements

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Lightware to showcase breakthrough Google Meet and TPN MM...

Lightware will exhibit several major product innovations at ISE 2026, including the new USB-C BOOSTER-V1, Google Meet. integration for various Taurus UCX models...

23/12/2025

Nielsen, Roku Expand Measurement Partnership

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

PwC: Streaming Market Shifting to 'Scale and Sustainability'

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Inside the Gray Innovation Lab

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

ESPN Renews Deal for Heisman Trophy Coverage

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Gray Media to Acquire WBBJ from Bahakel Communications

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

23/12/2025

Taking the Stage at Carnegie Hall-On a Global Scale

Taking the Stage at Carnegie Hall-On a Global Scale Boston Conservatory Orchestra students reflect on their epic concert marking the 80th session of the UN Gene...

23/12/2025

Boost Performance with a System Effectiveness Review

Experience the power of WO Automation for Radio's newest service, the System Effectiveness Review. Designed to help you achieve more, a System Effectiveness...

23/12/2025

How Steamy Can It Get? Single's Inferno' Season 5 Premieres January 20, Previews All-Out Flirting War in Sizzling Teaser

Back to All News How Steamy Can It Get? Single's Inferno' Season 5 Pre...

23/12/2025

33 Million Global Viewers on Netflix Watched Jake Paul vs. Anthony Joshua's Epic Six-Round Battle

Back to All News 33 Million Global Viewers on Netflix Watched Jake Paul vs. Ant...

23/12/2025

December 22, 2025

New technique lights up where drugs go in the body, cell by cell Scripps Research scientists developed a technique that maps drug binding in individual cells th...

22/12/2025

SVG New Sponsor Spotlight: Presidio's Neerav Shah on the Role of Its Captivate and Resonate Platforms in Sports Production

SVG New Sponsor Spotlight: Presidio's Neerav Shah on the Role of Its Captiva...

22/12/2025

Hitting the Bullseye: Sky Sports Readies Itself for the Biggest PDC World Darts Championship to Hit Ally Pally Yet

Hitting the bullseye: Sky Sports readies itself for the biggest PDC World Darts ...

22/12/2025

Unique Skillset: Bringing New Directors to the World of Darts at The Worlds with Sky Sports

Unique skillset: Bringing new directors to the world of darts at The Worlds with...

22/12/2025

Gravity Media Prepares for a Flight of Fancy With the PDC World Darts Championship 2025 for Sky Sports

Gravity Media prepares for a flight of fancy with the PDC World Darts Championsh...

22/12/2025

One Hundred and Eighty: Gravity Media on Hitting the Production Bullseye at the World Darts Championship 2025

One hundred and eighty: Gravity Media on hitting the production bullseye at the ...

22/12/2025

The Famous Group's Jon Slusser on Fascinating Fans Through Immersive Content Experiences

The Famous Group's Jon Slusser on Fascinating Fans Through Immersive Content...

22/12/2025

ESPN's Meg Aronowitz on Continuing High-Quality Broadcasts of Collegiate Sports, Expanding Growth of Internal Production Team

ESPN's Meg Aronowitz on Continuing High-Quality Broadcasts of Collegiate Spo...

22/12/2025

ESPN Takes Data-Driven Storytelling to New Heights with MNF Playbook with Next Gen Stats' NFL Altcasts

ESPN Takes Data-Driven Storytelling to New Heights with MNF Playbook with Next ...

22/12/2025

Paramount and Netflix Boast Double-Digit Gains in Nielsen's November Media Distributor Gauge

Paramount Scores Largest Share Increase Among Distributors as Paramount and CBS...

22/12/2025

Nielsen and Roku Expand Strategic Measurement Partnership

New multi-year deal integrates Roku's data to fuel Nielsen's measurement suite Roku gains access to Nielsen's streaming ratings, showing The Roku C...

22/12/2025

Allen Media Group to Deploy Infillion TrueX for Streaming Services

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

22/12/2025

Berklee Wrapped 2025: Our Top News and Stories

Berklee Wrapped 2025: Our Top News and Stories A look back at a year highlighted by faculty milestones, major film and television projects, Bob Dylan's ho...

22/12/2025

Marine Biological Laboratory Explores Human Memory With AI and Virtual Reality

The works of Plato state that when humans have an experience, some level of change occurs in their brain, which is powered by memory - specifically long-term me...

22/12/2025

Simplify Playlist Management with Workflows in WO Automation for Radio

Workflows allow you to create a sequence of planned events which may be added to your template(s) or inserted directly into your sequential or background playli...

22/12/2025

Global Anime Hits and New Releases Take Center Stage at Jump Festa 2026

Back to All News Global Anime Hits and New Releases Take Center Stage at Jump Festa 2026 Entertainment 22 December 2025 GlobalJapan Link copied to clipboar...

21/12/2025

Legoshi and Haru's Story Reaches Its Finale: BEASTARS Final Season Part 2' Premieres March 2026, Main Trailer Out Now

Back to All News Legoshi and Haru's Story Reaches Its Finale: BEASTARS Fin...

20/12/2025

Atomos Updates Ninja TX GO-Ninja TX With ProRes RAW and C...

Atomos announced the immediate availability of a new firmware update for its Ninja TX GO and Ninja TX monitor-recorders, unlocking ProRes RAW recording from the...

20/12/2025

CJP Broadcast Completes Digitisation of European Gymnasti...

CJP Broadcast has completed the digitisation of the European Gymnastics tape archive, converting 328 tapes containing more than forty years of recorded material...

20/12/2025

Bitmovin Launches Stream Lab MCP Server

Bitmovin, the leading provider of video streaming solutions, today announced the launch of the Stream Lab MCP Server, to give AI agents and large language model...

20/12/2025

Gracenote Unveils New Immersive Features for Sports Hubs

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

20/12/2025

Morgan Murphy Media Promotes Jill Shiroma to VP of Digital Strategy

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

20/12/2025

Samsung Makes GameBreaks Ad Format Available Programmatically

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

20/12/2025

FCC Extends Deadline for Comments on Upper C-Band Proposals

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

20/12/2025

Scripps Sports Inks Deals for Soccer and Professional Cheerleading

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

20/12/2025

Atomos Unveils Firmware Update For Ninja TX and TX GO

Share Share by: Copy link Facebook X Whatsapp Pinterest Flipboard...

20/12/2025

Barack Obama Includes Laufey on His 2025 Favorite Music List

Barack Obama Includes Laufey on His 2025 Favorite Music List The former presidents roundup of books, music, and movies includes a song from the Berklee alums ...

20/12/2025

December 19, 2025

Study reveals a key hormonal circuit in the kidneys Scripps Research scientists identify the protein that helps kidney cells regulate renin, providing foundatio...

19/12/2025

SVG Sit-Down: Diversified's Jared Timmins on AI for Broadcast Sports and Creating the Smart Venue'

SVG Sit-Down: Diversified's Jared Timmins on AI for Broadcast Sports and Cre...

19/12/2025

2025 SVG Summit Audio Recap: Say What?

2025 SVG Summit Audio Recap: Say What?The Audio Production and Distribution Workshop at the SVG Summit 20 took on issues including speech intelligibility, Next-...