
NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.
More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.
NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.
NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.
Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.
Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.
At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.
Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.
Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.
JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.
JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.
https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4
New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.
The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.
FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.
NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.
Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.
VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.
The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.
VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
07/10/2026
Dalet, a leading technology and service provider for media-rich organizations, today announced the latest Long-Term Supported (LTS) release of Dalet Flex. Build...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
27/07/2026
The Audio Engineering Society Educational Foundation has named its scholarship and grant recipients for the 2026/2027 academic year. The awards support students...
27/07/2026
ESPN and the Mid-Eastern Athletic Conference (MEAC) have signed a multi-year media rights agreement covering basketball championships and select football and ba...
27/07/2026
Yahoo Sports and Cameo have partnered to allow fantasy football users to book personalized celebrity videos directly within the Yahoo Fantasy app and website. T...
27/07/2026
Eutelsat has welcomed the FCC's order establishing the regulatory framework for the transition of Upper C-band spectrum in the United States. Under the orde...
27/07/2026
ESPN and the International Dance League (IDL) have signed a global rights agreement for the 2026 season. Events will air bi-weekly on ESPN2 at 7 p.m. ET beginni...
27/07/2026
NBCUniversal and YouTube have signed a multi-year global partnership that includes bundling Peacock with YouTube Premium, extending NBCUniversal's distribut...
27/07/2026
Agenda includes presentations from the Miami Heat, Inter Miami, Florida Panthers...
27/07/2026
Behind The Mic provides a roundup of recent news regarding on-air talent, including new deals, departures, and assignments compiled from press releases and repo...
27/07/2026
Ten one-minute shorts will roll out across CBS, The CW, Paramount+ during the 2026 PBR Teams season
PBR is officially moving into scripted animation. The Profe...
27/07/2026
Today the nonprofit Sundance Institute shared the names of this year's cohor...
27/07/2026
New app turns vocal percussion into polished loops
Vochlea, the company behind the voice-to-MIDI instrument Dubler 2, have just announced the launch of a ne...
27/07/2026
Latest MPE-capable instrument completes Pure Trilogy line-up
Sonora Cinematic have just released the third and final instalment in their series of acoustic ...
27/07/2026
Emulates four varieties of optical compression
The latest addition to the fedDSP line-up has just been announced, and packs four distinct optical compressor...
27/07/2026
Optical & FET compressors join 500-series line-up
At our recent GearExpo UK show, we heard from Drawmer that they would soon be launching a pair of new 500-...
27/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
27/07/2026
Lightware announces SNMPv3-based monitoring support for selected Taurus devices, available from firmware v1.20.0. The update enables IT administrators and AV in...
27/07/2026
Mediagenix, a global leader in smart content solutions to profitably connect the right content to the right audience, will showcase its vision for the Real-Time...
27/07/2026
All episodes available Thursday 6 August on Sky and streaming service NOWMonday ...
27/07/2026
Why Confidence MattersIn this series, we explore the technologies, architectures and operational realities shaping modern media operations. Along the way, we ex...
27/07/2026
Apple today announced AppleCare One - a new way for customers to protect multipl...
27/07/2026
Yesterday witnessed truly memorable scenes in Croke Park, bringing the curtain d...
27/07/2026
Luxembourg, July 27, 2026 - SES, a leading space solutions company, issued the following statement following the U.S. Federal Communications Commission approval...
27/07/2026
Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government...
26/07/2026
What is Audio Vivid? A New Destination for Immersive Audio Over the past several years, immersive audio has evolved from a niche production consideration into a...
26/07/2026
The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems. To help meet t...
25/07/2026
New app turns vocal percussion into polished loops
Vochlea, the company behind the voice-to-MIDI instrument Dubler 2, have just announced the launch of a ne...
25/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/07/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
25/07/2026
Export FCP7 to XML Without Needing FCP7 Installed
Dave Rice July 24, 2026
0 Comments
I've created an open source tool that allows you to export FC...
25/07/2026
Documentary Sell Your House Edited with DaVinci Resolve Studio
Brie Clayton July 24, 2026
0 Comments
Filmmakers post feature documentary from opposite...
25/07/2026
Time to Act: Bringing a New Opera to Life Boston Conservatory at Berklee vocal arts students played a vital role in shaping the opera's creative developme...
25/07/2026
Beyond the Score For Dean of Music Kevin Haden, the boundless possibilities of classical music begin with process.
July 24, 2026
By
Kevin Haden
Michelle ...
25/07/2026
Reading the Room As music director of two major orchestras, conductor Jonathon Heyward aims to make the concert hall a space for everyone.
July 24, 2026
By
...
25/07/2026
Saturday 25 July 2026
First look revealed for YAGA, a seductive mystery thrille...
24/07/2026
Referee hat cameras, player/coach mics, and virtual graphics will be deployed at the three-day event for the first time
The NFL FLAG Championships return to ES...
24/07/2026
Qvest, Ateme, and Scaleway have partnered to offer broadcasters, media organizat...
24/07/2026
LTN has completed satellite-to-IP migrations for more than 3,000 broadcaster, MVPD (Multi-Channel Video Programming Distributor), head-end, and content owner si...
24/07/2026
Chyron has released Virtual Placement 8.1, an update to its live sports virtual graphics platform. The release adds workflow and graphics changes across chroma ...
24/07/2026
Riedel's RefCam Live was used for the first time at an international soccer match during a friendly between the German men's national team and Ghana in ...
24/07/2026
USSI Global said it is prepared to help broadcasters, satellite network operators, and media organizations manage the next phase of C-band spectrum clearing, fo...
24/07/2026
FuboTV and The Athletic have launched The Athletic Video Hub on Fubo, bringing The Athletic's sports video content to connected TV (CTV) for the first time....