Sony Pixel Power calrec Sony

Seamless in Seattle: NVIDIA Research Showcases Advancements in Visual Generative AI at CVPR

17/06/2024

NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.

More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.

NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.

NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.

Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.

Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.

At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.

Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.

Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.

JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.

JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.

https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4

New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.

The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.

FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.

NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.

Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.

VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.

The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.

VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
LINK: https://blogs.nvidia.com/blog/visual-generative-ai-cvpr-research/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

03/04/2026

2026 NAB Show Exhibitor Insight: Techex

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

TelevisaUnivision Signs New Nielsen Media Intelligence Deal

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

The WNET Group, JIB Launch NHK World-Japan in New York

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

NAB Leadership Foundation Welcomes New Board Members

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

EverPass Media Expands Distribution Deal with Netflix

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

Versant Acquires AI-Data Platform StockStory

Share Copy link Facebook X Linkedin Bluesky Email...

03/04/2026

CVP Grows European Footprint with Strategic Expansion in...

CVP, one of Europe's leading suppliers of professional video and broadcast solutions, today announces the launch of its new German operation and the formati...

03/04/2026

MRMC Announces Appointment of Chief Operating Officer

Mark Roberts Motion Control (MRMC) today announces the appointment of Nick Barthee as Chief Operating Officer, strengthening its leadership as the company conti...

03/04/2026

Net Insight Introduces Programmable Trust Boundaries for...

Net Insight introduces programmable Trust Boundaries that make live media interconnection predictable as traffic moves between facilities, networks and cloud en...

03/04/2026

Winning in the new media economy: Avid showcases AI-powered, connected intelligence to unlock media value at NAB Show 2026

Winning in the new media economy: Avid showcases AI-powered, connected intellige...

03/04/2026

NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibility and Loudness at NAB 2026

NUGEN Audio CEO Dr. Paul Tapper to Lead Presentation About Dialog Intelligibilit...

03/04/2026

NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances

NAB Show 2026: PlayBox Neo Highlights Workflow, Security, and IP Advances Brie Clayton April 2, 2026 0 Comments PlayBox Neo will showcase the latest i...

03/04/2026

For Taku Hirano, Everything Is Connected

For Taku Hirano, Everything Is Connected From touring and composition to teaching and instrument design, the in-demand percussionist sees it all as one body o...

03/04/2026

Berklee Honors Humberto Ramirez with Master of Latin Music Award

Berklee Honors Humberto Ramirez with Master of Latin Music Award The alumnus and acclaimed trumpeter is honored for his influence as a performer, composer, an...

02/04/2026

HBO and NFL Films Announce Hard Knocks: Training Camp with the Seattle Seahawks, Debuting August 11

HBO and NFL Films have announced Hard Knocks: Training Camp with the Seattle Sea...

02/04/2026

NAB 2026: Haivision Unveils Makito ONE Video Transport Platform

Haivision has announced the Makito ONE, a single-blade video encoding and decoding platform, at NAB Show 2026. The platform combines dual-channel video encoding...

02/04/2026

NAB 2026: Telestream Introduces UP.Lens Cloud-Based Multiviewer and Monitoring Service

Telestream has introduced UP.Lens, a cloud-based multiviewer and monitoring serv...

02/04/2026

NAB 2026: MRMC to Showcase Robotic Camera Technology and Mark 60th Anniversary

Mark Roberts Motion Control (MRMC) will exhibit at NAB Show 2026 (Booth C5220, April 19-22, Las Vegas Convention Center), marking the company's 60th anniver...

02/04/2026

NAB 2026: Net Insight Introduces Programmable Trust Boundaries for Live Media Interconnection

Net Insight has introduced programmable Trust Boundaries, a feature integrated i...

02/04/2026

NAB 2026: Bitmovin Adds SGAI Support to Playback Products

Bitmovin has announced support for SGAI (Server-Guided Ad Insertion) in its playback products, using HLS interstitials. SGAI combines elements of client-side an...

02/04/2026

Binghamton University Athletics Adds Riedel SimplyLive RiMotion R12 for Student-Run Productions

Riedel Communications' SimplyLive RiMotion R12 replay system is supporting B...

02/04/2026

NAB 2026: LTN and Ateme Announce Integration of Video Processing with IP Transport

LTN, a managed IP video transport company, and Ateme, a video compression and de...

02/04/2026

TDF Expands Channel Capacity on Terrestrial Broadcast Network with Harmonic

Harmonic has announced that TDF, a broadcast infrastructure operator in France, has deployed Harmonic's XOS Advanced Media Processor and ProStream X Video S...

02/04/2026

United Rugby Championship Reports First-Year Results with Eluvio Streaming Platform

Eluvio and the United Rugby Championship (URC) have announced first-year results...

02/04/2026

Kansas City Current and Scripps Sports Announce ION as Broadcast Home of 2026 Teal Rising Cup

The Kansas City Current and Scripps Sports have announced that ION will broadcas...

02/04/2026

ESPN Announces Courtside Alt-Cast for Women's Final Four

ESPN will debut Courtside at the Women's Final Four Presented by AT&T, an alt-cast airing Friday, April 3 at 7 p.m. and 9:30 p.m. ET on ESPN2, and Sunday, A...

02/04/2026

PAMA and Shure Accept Applications for 6th Annual Mark Brunner Professional Audio Scholarship

The Professional Audio Manufacturers Alliance (PAMA) and Shure Incorporated are ...

02/04/2026

ESPN's MegaCast Coverage of 2026 NCAA Women's Final Four Begins Friday, April 3 in Phoenix

ESPN's MegaCast Coverage of 2026 NCAA Women's Final Four Begins Friday, ...

02/04/2026

DAZN Launches Playmakers Creator Program

DAZN has announced the launch of DAZN Playmakers, a global influencer program designed to build a network of sports content creators. The programme will give cr...

02/04/2026

NFL Network Unveils New Production Ops Leadership Structure as ESPN Takes Over

Tony Cole, Jessica Lee shift into new roles reporting to ESPN SVP/Content Operations Chris Calcinari....

02/04/2026

TNT Sports and CBS Sports to Reunite Michigan's Iconic Fab Five for Special NCAA Men's Final Four Altcast on truTV & HBO Max

Michigan's Fab Five will reunite for an alternate presentation of the Mich...

02/04/2026

Coming of Age: ESPN and NHL's Inside Out Classic Marks Another Step Forward for the Animated Alternative Broadcast

Real-time tracking, virtual production, and Pixar storytelling converge for Apri...

02/04/2026

Streaming Around the Moon: NASA+ Goes Live From Historic Artemis II Mission

NASA's long-awaited Artemis II mission has launched four astronauts on a 10-day journey around the moon, marking the first manned launch toward the moon sin...

02/04/2026

Release Rundown: What to Watch in April, From Bunnylovr to Omaha

(L-R) Molly Belle Wright, Wyatt Solis, and John Magaro appear in Omaha by Cole Webley, an official selection of the 2025 Sundance Film Festival. (Photo courte...

02/04/2026

VEMIA Auction incoming

4 - 11 April 2026 VEMIA's latest gear auction is just around the corner, and there's already a wealth of sought-after instruments and studio gear li...

02/04/2026

Universal Audio's Voice Of God now native

Little Labs emulation now available to all The latest Universal Audio plug-in to become available outside of the company's DSP-powered UAD2 platform has...

02/04/2026

SBS Welcomes Back Cup Fever! For The Biggest FIFA World Cup Ever

SBS Welcomes Back Cup Fever! For The Biggest FIFA World Cup Ever 2 April, 2026 Media releases Santo Cilauro, Ed Kavalee and a host of special guests retur...

02/04/2026

Truth. Power. Perspective: NITV's Flagship Current Affairs Programs Return to Lead the National Conversation

Truth. Power. Perspective: NITV's Flagship Current Affairs Programs Return t...

02/04/2026

L3Harris Powers First Crewed Mission Around the Moon in 50 Years

L3Harris has successfully powered the historic launch of the Artemis II mission, providing propulsion and avionics....

02/04/2026

Scripps Completes Sale of WRTV to Circle City Broadcasting

Share Copy link Facebook X Linkedin Bluesky Email...

02/04/2026

GoVertical! AiDi Powers Real-Time 9:16 Autocropping for I...

Already deployed extensively by NBC Sports, FOR-A Corporation will demonstrate GoVertical! AiDi, the real-time 9:16 autocropping feature of viztrick AiDi, durin...

02/04/2026

Elite Media Technologies Selects Interra Systems BATON Fi...

Interra Systems, a provider of end-to-end quality assurance solutions for the digital media industry, announced that Elite Media Technologies has selected its B...

02/04/2026

TDF Expands Broadcast Channel Lineup with Harmonic

Harmonic's Media Processing Solutions Maximize Bandwidth Efficiency for Terrestrial Broadcast Delivery Harmonic (NASDAQ: HLIT) today announced that TDF, a...

02/04/2026

FOR-A's Software-Defined, AI-Powered Development Advances...

NBC Sports Deploys viztrick AiDi to Stream Live Events in 9:16 Mobile-First Formats with Auto Tracking, Development Signals Strategic Shift for FOR-A Long reco...

02/04/2026

Evergent showcases innovations in sports streaming and mo...

Evergent will showcase new innovations in subscriber lifecycle management and monetization at NAB Show 2026 (Las Vegas, April 18 22), including: New advances i...

02/04/2026

Binghamton University Strengthens Student Run Productions...

Riedel Communications is proud to be part of Binghamton University, State University of New York, Athletics' milestone year, celebrating the university'...