Sony Pixel Power calrec Sony

Seamless in Seattle: NVIDIA Research Showcases Advancements in Visual Generative AI at CVPR

17/06/2024

NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.

More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.

NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.

NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.

Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.

Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.

At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.

Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.

Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.

JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.

JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.

https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4

New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.

The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.

FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.

NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.

Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.

VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.

The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.

VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
LINK: https://blogs.nvidia.com/blog/visual-generative-ai-cvpr-research/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

01/04/2026

DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION

January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION Douyin Users Can Now Create And Share Videos With Stun...

05/03/2026

Clear-Com Supplies Cloud-Based Communications System for...

Clear-Com has provided Gen-IC virtual intercom, its cloud-based voice communications system for SaxaVord Spaceport, the first fully licensed vertical launch s...

05/03/2026

Hollywood Professional Association Concludes the 2026 HPA...

The Hollywood Professional Association (HPA) concluded the 2026 HPA Tech Retreat, convening more than 800 industry leaders, technologists, creatives and executi...

05/03/2026

LynTec 20 AMP Dual Output DMX Relay Now Shipping

LynTec, a leading manufacturer of electrical power control solutions for professional audio, video, and lighting systems, today announced that its new Dual DMX ...

05/03/2026

Snicket Labs Partners with ES Broadcast to Expand Access...

Snicket Labs is pleased to announce a new distribution partnership with ES Broadcast for its award winning solutions, Match and Enrich. Under the agreement, ES...

05/03/2026

LiveU Marks First Large Scale Global Deployment of AI Dri...

LiveU today announced the first large-scale deployment of its AI-driven LiveU IQ (LIQ ) technology at a global, multi-venue sporting event, setting a new benchm...

04/03/2026

Lega Basket Serie A Modernizes Media Operations Across Italian Basketball with ScorePlay

Lega Basket Serie A (LBA), the governing body for Italy's premier basketball...

04/03/2026

To Mark Five Years at Wrexham AFC, Co-Chairmen Rob Mac and Ryan Reynolds Will Host Broadcast During Wrexham-Swansea City on March 13

Wrexham AFC co-chairmen Rob Mac and Ryan Reynolds will host a first-of-its-kind ...

04/03/2026

FOX Sports Marks 100 Days to FIFA World Cup 2026 with Company-Wide Celebration

The countdown is underway and with just 100 days to go until the world's greatest sporting event begins on Thurs., June 11, FOX Sports, America's Englis...

04/03/2026

Exchange, NBCUniversal Team Up to Provide Service Members with Free Streaming of Paralympic Winter Games

No matter where they are in the world, service members and veterans can stream N...

04/03/2026

Telemundo Releases Somos Ms, the Official Anthem of its FIFA World Cup 2026 Coverage

Telemundo officially releases Somos M s, the anthem for the network's cove...

04/03/2026

Hollywood Professional Association Concludes 2026 HPA Tech Retreat

The Hollywood Professional Association (HPA) concluded the 2026 HPA Tech Retreat, convening more than 800 industry leaders, technologists, creatives, and execut...

04/03/2026

Case Study: How SEG+ Unified Utah's Biggest Sports Teams into a 40%+ Subscriber Growth Streaming Platform

Smith Entertainment Group transformed how local sports are consumed by creating ...

04/03/2026

SVG in Indy: Pacers Sports & Entertainment Remotely Produces Broadcasts of G League's Noblesville Boom

The production method, which spans a distance of 36.2 miles, was designed and im...

04/03/2026

Haivision Releases Seventh Annual Broadcast Transformation Report, Highlights Key Trends Shaping Live Production in 2026

Haivision, a global provider of mission-critical, real-time video networking and...

04/03/2026

SVG Sit-Down: Quantum CEO Hugues Meyrath on Reshaping the Company, the Impact of AI, Evolving Media Storage

When Hugues Meyrath came out of retirement to take the helm as CEO of Quantum, i...

04/03/2026

As Banana Ball Expands Exponentially, So Too Do Its Production Capabilities

Last week's launch of Banana Ball Championship League has spurred a significant upgrade of production facilities The Savannah Bananas, arguably the hottest...

04/03/2026

Press Release TEST

sldkfjsdlfkjsldkfjsldkjfslkdjfslkdjfsl Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus...

04/03/2026

Spotify A/Presenta Brings Fans Closer to Artists' Creative Process in Latin America, Starting With ROSALA

For many fans, a song's backstory can be just as compelling as the final pro...

04/03/2026

Spotify Kicks Off Our 20th Anniversary at SXSW With a Celebration of Artists, Creators, and Fans

In 2006, Spotify was founded on the belief that technology could bring artists a...

04/03/2026

Spotify and Coca-Cola Saddle Up for a Rhinestone Cowboy Experience at the Houston Rodeo

Spotify is back on the ground for the Houston Livestock Show and Rodeo, and we&#...

04/03/2026

Spotify Doubles Down on Investments in Australian Fan Discovery of Homegrown Aussie Talent

Spotify had an energizing week in Sydney, Australia, filled with powerful conver...

04/03/2026

FKA twigs and Jordan Hemingway Explore the Making of HARD' in Episode Two of Directed By'

Earlier this year, we launched Directed By, a documentary-style series that pull...

04/03/2026

100 Days to Go: SBS Unveils World Class Team for The Greatest Show on Earth -The FIFA World Cup 2026

100 Days to Go: SBS Unveils World Class Team for The Greatest Show on Earth -The...

04/03/2026

KT and Rohde & Schwarz to showcase AI-enhanced radio transmission performance

KT and Rohde & Schwarz to showcase AI-enhanced radio transmission performance In a joint 6G AI proof-of-concept demonstration, the CMX500 one-box tester from ...

04/03/2026

APTS Announces Public Broadcast Leadership, Advocacy Awards

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

SES publishes 2025 Annual Report

Luxembourg, 3 March 2026 - SES S.A. has today published its 2025 Annual Report, following the announcement of the company's full year financial results for ...

04/03/2026

NBC Sports, USA Sports Extend Rights Deal with PGA

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

Broadcasters Gather in DC for NAB State Leadership Conference

Share Copy link Facebook X Linkedin Bluesky Email...

04/03/2026

Transforming Africa's Future Farmers: Satellite-Enabled IoT Powers Data-Driven Agribusinesses

Luxembourg, March 3, 2026 - SES, a leading space solutions company, along with I...

04/03/2026

VEON Partners with GSMA Innovation Fund to Accelerate Digital Innovation in Pakistan and Bangladesh

04 Mar 2026 VEON Partners with GSMA Innovation Fund to Accelerate Digital Innov...

04/03/2026

VEON's Beeline Uzbekistan and Rakuten Symphony Partner for Open RAN, AI Collaboration

04 Mar 2026 VEON's Beeline Uzbekistan and Rakuten Symphony Partner for Open...

04/03/2026

Sky Sports unveils plans for 2026 Formula 1 coverage

Wednesday 4 March 2026 Sky Sports unveils plans for 2026 Formula 1 coverage Sky Sports is preparing for one of the most highly anticipated F1 seasons in recen...

04/03/2026

Bloodhounds' Season 2 Gears Up for April 3 Premiere with Hard-Hitting New Teaser and Poster

Back to All News Bloodhounds' Season 2 Gears Up for April 3 Premiere with ...

04/03/2026

Netflix Ads Suite Expands Capabilities

Back to All News Netflix Ads Suite Expands Capabilities Business 04 March 2026 GlobalUnited States Link copied to clipboard After launching the Netflix Ad...

04/03/2026

March 02, 2026

Scripps Research welcomes healthcare innovator Joe Kiani to the Board of Directors Kiani brings decades of experience in patient safety and public service. Mar...

04/03/2026

March 03, 2026

Nanoparticle vaccine approach takes on a new target: Hepatitis C virus Scripps Research scientists reengineer critical proteins on the surface of HCV, paving th...

03/03/2026

LIV Golf, Beyond Sports Elevate Online Gaming Ecosystem with Launch of LIV Golf Fantasy and LIV X

Beyond Sports, a Sony group company, and LIV Golf, the world's golf league, ...

03/03/2026

Ilitch Sports + Entertainment Announces Launch of Detroit SportsNet

Ilitch Sports + Entertainment announces the launch of Detroit SportsNet (DSN), a year-round broadcast home for two of Detroit's franchises. With flexible op...

03/03/2026

Advanced Systems Group Promotes Gretchen Taipale to Vice President, Managed Services

Advanced Systems Group, LLC (ASG), a technology and services provider for media ...

03/03/2026

PGA of America, NBC Sports, and USA Sports Extend Media Rights Agreement Through 2033

The PGA of America, NBC Sports and USA Sports extend their media rights agreemen...

03/03/2026

HONOR, ARRI Announce Technical Collaboration to Bring ARRI Image Science into Next-Gen Consumer Devices

AI device ecosystem company HONOR enters into a strategic technical collaboratio...

03/03/2026

Telos Alliance Partners with College Radio Foundation to Support College Broadcasters

Cleveland's Telos Alliance, pioneers in broadcast technology for 30 years, l...

03/03/2026

Sennheiser Relaunches MD 9235 Wireless Mic Head

The MD 9235 microphone head for wireless handhelds has been a firm favorite with many engineers and artists for its ability to cut through high on-stage levels ...

03/03/2026

Haivision to Showcase Private 5G and Live Video Contribution Innovations at MWC 2026

Haivision Systems Inc. (Haivision), a global provider of mission-critical, real-...