
NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.
More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.
NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.
NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.
Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.
Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.
At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.
Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.
Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.
JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.
JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.
https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4
New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.
The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.
FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.
NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.
Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.
VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.
The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.
VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
North America Stories
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/10/2026
Cobalt Digital to Bring End-to-End ST 2110/IPMX Solutions and Practical Tools to NAB Show New York
Highlights include complete path from SDI to IP, C-band tran...
07/10/2026
Deerstalker Pictures The Exorcism of Nixie Shot With Blackmagic Cameras
Brie Clayton October 6, 2026
0 Comments
Fantasy series relied on Blackmagic ca...
07/10/2026
Nevion announces Providius NVRT integration with eMerge switches for expanded br...
07/10/2026
DriveShelf - A Filmmaker's Mac App for Finding Footage on Unplugged Drives
Brie Clayton October 6, 2026
0 Comments
DriveShelf for Mac catalogs eve...
07/10/2026
COW Jobs: Cinematographer, Video Editor - Baltimore, MD
Brie Clayton October 7, 2026
0 Comments
Cinematographer / Video Editor
September 23, 2026 ...
07/10/2026
Behind the Scenes at Berklee Homecoming 2026 Follow Marley Striem BM '25 behind the scenes as she manages a team of students and liaises with artists duri...
07/10/2026
Berklee College of Music and Berklee Valencia Named to Billboards 2026 List of T...
06/10/2026
A response to Mark Turner and the SVG AI Innovation Lab
SVG AI Innovation Lab Program Director Mark Turner's How Broadcast Sports Can Apply the Newsroom S...
06/10/2026
Cobalt Digital will exhibit at NAB Show New York 2026 (Booth 226, Javits Center, October 21-22), demonstrating its ST 2110/IPMX product line alongside solutions...
06/10/2026
LiveU and Airwise have announced a strategic partnership integrating LiveU's bonded video transmission with Airwise's drone fleet management, airspace a...
06/10/2026
Hotspur Labs, Tottenham Hotspur's corporate venture programme, will co-host ...
06/10/2026
SPORTEL Monaco 2026 will take place October 19-21 at the Grimaldi Forum in Monaco, with the conference programme running October 19-20.
Sessions include:
Mast...
06/10/2026
Zixi has announced the appointment of John Towers as Chief Financial Officer. Towers brings more than 20 years of experience in media and entertainment, includi...
06/10/2026
The Memphis Grizzlies, DAZN, and WATN/ABC24 have announced a partnership to simulcast 15 regular season Grizzlies games free over-the-air on ABC24 during the 20...
06/10/2026
The Minnesota Timberwolves, KARE 11, and DAZN have announced a partnership to simulcast 15 Timberwolves games free over-the-air on KARE 11 during the 2026-27 se...
06/10/2026
Sportel 2026 is set to be held in Monaco in less than two weeks (beginning Octob...
06/10/2026
The company's Pro Data portable storage unit enabled the Dallas Cowboys prod...
06/10/2026
Manufacturer receives honors from both TV Tech and TVB Europe for advancements in compression technology
AMSTERDAM September 30, 2026 - Cobalt Digital has ...
06/10/2026
Cobalt Digital NAB NY Booth 226 // Journalists: Click to visit Cobalt
Highlight...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Imaginary Forces Paints a Portrait of Conflict in War
Brie Clayton October 6, 2026
0 Comments
Director Ronnie Koff of Imaginary Forces (IF) transforms...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Pebble, a leading specialist in automation, content management and integrated channel solutions, is supporting the continued modernisation of playout at Brazili...
06/10/2026
Techex, a specialist in software-defined live video workflows, will demonstrate the how tx darwin platform combines satellite and IP delivery at NAB New York 20...
06/10/2026
NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
Allen Bourgoyne October 5, 2026
0 Comments
Local AI is becoming more usef...
06/10/2026
Cannes Premiere Documentary The Match Shot with Blackmagic URSA Cine
Brie Clayton October 5, 2026
0 Comments
Portraiture and archival footage bring pl...
06/10/2026
Boston Conservatory Artists Turn Chairs on The Voice Watch Lexi Stephens (BFA 25) and Marcus Ladden Jr. (BFA 30) wow the celebrity coaches during their blind au...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/10/2026
Sean Giancola to Step Down as Publisher & CEO of New York Post Media Group
New York (October 6, 2026) - After more than a decade of transformative leadership a...
06/10/2026
Telecom operators are increasingly building their AI strategies on open models - and the reasons go beyond mere cost.
Open models give telcos the ability to t...
05/10/2026
Eutelsat has announced that Airbus Defence and Space has completed the first 32 new OneWeb Low Earth Orbit satellites at its facility in Toulouse, France. The s...
05/10/2026
Racecourse Media Group (RMG) has appointed Andrew Demaria as its first Chief Content Officer, effective October 6.
Demaria joins from a career spanning more th...
05/10/2026
SMPTE will host SMPTE Bootcamp: Precision Timing for ST 2110 on Wednesday, October 21, 2026, from 9:20 a.m. to 4:30 p.m. in Room 3D03 at the Javits Center durin...
05/10/2026
Nexstar Media Group and DAZN have announced a partnership to simulcast 15 regular season Indiana Pacers games on Nexstar television stations and free on DAZN si...
05/10/2026
Behind The Mic provides a roundup of recent news regarding on-air talent, including new deals, departures, and assignments compiled from press releases and repo...
05/10/2026
Telos Alliance has announced that Zephyr Connect SE, a new hardware codec, is now shipping. The unit joins the Telos iPort High Density multi-codec gateway and ...