Sony Pixel Power calrec Sony

Seamless in Seattle: NVIDIA Research Showcases Advancements in Visual Generative AI at CVPR

17/06/2024

NVIDIA researchers are at the forefront of the rapidly advancing field of visual generative AI, developing new techniques to create and interpret images, videos and 3D environments.

More than 50 of these projects will be showcased at the Computer Vision and Pattern Recognition (CVPR) conference, taking place June 17-21 in Seattle. Two of the papers - one on the training dynamics of diffusion models and another on high-definition maps for autonomous vehicles - are finalists for CVPR's Best Paper Awards.

NVIDIA is also the winner of the CVPR Autonomous Grand Challenge's End-to-End Driving at Scale track - a significant milestone that demonstrates the company's use of generative AI for comprehensive self-driving models. The winning submission, which outperformed more than 450 entries worldwide, also received CVPR's Innovation Award.

NVIDIA's research at CVPR includes a text-to-image model that can be easily customized to depict a specific object or character, a new model for object pose estimation, a technique to edit neural radiance fields (NeRFs) and a visual language model that can understand memes. Additional papers introduce domain-specific innovations for industries including automotive, healthcare and robotics.

Collectively, the work introduces powerful AI models that could enable creators to more quickly bring their artistic visions to life, accelerate the training of autonomous robots for manufacturing, and support healthcare professionals by helping process radiology reports.

Artificial intelligence, and generative AI in particular, represents a pivotal technological advancement, said Jan Kautz, vice president of learning and perception research at NVIDIA. At CVPR, NVIDIA Research is sharing how we're pushing the boundaries of what's possible - from powerful image generation models that could supercharge professional creators to autonomous driving software that could help enable next-generation self-driving cars.

At CVPR, NVIDIA also announced NVIDIA Omniverse Cloud Sensor RTX, a set of microservices that enable physically accurate sensor simulation to accelerate the development of fully autonomous machines of every kind.

Forget Fine-Tuning: JeDi Simplifies Custom Image Generation Creators harnessing diffusion models, the most popular method for generating images based on text prompts, often have a specific character or object in mind - they may, for example, be developing a storyboard around an animated mouse or brainstorming an ad campaign for a specific toy.

Prior research has enabled these creators to personalize the output of diffusion models to focus on a specific subject using fine-tuning - where a user trains the model on a custom dataset - but the process can be time-consuming and inaccessible for general users.

JeDi, a paper by researchers from Johns Hopkins University, Toyota Technological Institute at Chicago and NVIDIA, proposes a new technique that allows users to easily personalize the output of a diffusion model within a couple of seconds using reference images. The team found that the model achieves state-of-the-art quality, significantly outperforming existing fine-tuning-based and fine-tuning-free methods.

JeDi can also be combined with retrieval-augmented generation, or RAG, to generate visuals specific to a database, such as a brand's product catalog.

https://blogs.nvidia.com/wp-content/uploads/2024/06/JeDi-cow-sculpture.mp4

New Foundation Model Perfects the Pose NVIDIA researchers at CVPR are also presenting FoundationPose, a foundation model for object pose estimation and tracking that can be instantly applied to new objects during inference, without the need for fine-tuning.

The model, which set a new record on a popular benchmark for object pose estimation, uses either a small set of reference images or a 3D representation of an object to understand its shape. It can then identify and track how that object moves and rotates in 3D across a video, even in poor lighting conditions or complex scenes with visual obstructions.

FoundationPose could be used in industrial applications to help autonomous robots identify and track the objects they interact with. It could also be used in augmented reality applications where an AI model is used to overlay visuals on a live scene.

NeRFDeformer Transforms 3D Scenes With a Single Snapshot A NeRF is an AI model that can render a 3D scene based on a series of 2D images taken from different positions in the environment. In fields like robotics, NeRFs can be used to generate immersive 3D renders of complex real-world scenes, such as a cluttered room or a construction site. However, to make any changes, developers would need to manually define how the scene has transformed - or remake the NeRF entirely.

Researchers from the University of Illinois Urbana-Champaign and NVIDIA have simplified the process with NeRFDeformer. The method, being presented at CVPR, can successfully transform an existing NeRF using a single RGB-D image, which is a combination of a normal photo and a depth map that captures how far each object in a scene is from the camera.

VILA Visual Language Model Gets the Picture A CVPR research collaboration between NVIDIA and the Massachusetts Institute of Technology is advancing the state of the art for vision language models, which are generative AI models that can process videos, images and text.

The group developed VILA, a family of open-source visual language models that outperforms prior neural networks on key benchmarks that test how well AI models answer questions about images. VILA's unique pretraining process unlocked new model capabilities, including enhanced world knowledge, stronger in-context learning and the ability to reason across multiple images.

VILA can understand memes and reason based on multiple images or video frames. The VILA model fa
LINK: https://blogs.nvidia.com/blog/visual-generative-ai-cvpr-research/...
See more stories from nvidia

Most recent headlines

05/01/2027

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be demoed at CES 2026

Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...

01/06/2026

Dolby Sets the New Standard for Premium Entertainment at CES 2026

January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026 Throughout the week, Dolby brings to life the latest innovatio...

02/05/2026

Dalet Flex LTS Delivers Smarter Search, Faster Editing, and an AI-Ready Foundation for Modern Media

Dalet, a leading technology and service provider for media-rich organizations, t...

01/05/2026

NBCUniversal's Peacock to Be First Streamer to Integrate Dolby's Full Suite of Premium Picture and Sound Innovations

January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...

01/04/2026

DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION

January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION Douyin Users Can Now Create And Share Videos With Stun...

17/02/2026

Eutelsat, Viewsat Renew Capacity Agreements to Support Development of Broadcast Market in MENA Region

Eutelsat announces the renewal of multiple capacity agreements with Viewsat, rei...

17/02/2026

The Evolution of Communications in Cloud Era of Global Broadcasting

At Milano Cortina 2026, one statistic stood out across industry reporting: 70% of global signal distribution is now happening over public cloud. What began as e...

17/02/2026

CP Communications Supports the Big Game Broadcast for 27 Years

CP Communications provided broadcast support for Super Bowl LX in Santa Clara, Calif. at Levi's Stadium, marking 27 years of working around the Big Game. Th...

17/02/2026

NBC Sports Graphics Team Goes Onsite, Embraces L.A. Motif for NBA All-Star

Street-level images of iconic Los Angeles landmarks are incorporated into the regular-season look With NBC Sports just months into its return to NBA coverage a...

17/02/2026

Jimpa Asks You to Choose Love

Sophie Hyde and John Lithgow backstage during the premiere of Jimpa. (Photo by George Pimentel / Shutterstock for Sundance Film Festival)...

17/02/2026

The 19th Annual South African Film and Television Awards (SAFTAs19) Nominees Announced

Johannesburg, 17 February 2026 - The National Film and Video Foundation (NFVF), ...

17/02/2026

L3Harris Marks Delivery of 3 Million Fuzes to US Army

A Green Beret hangs a 60mm mortar round from a M225A mortar system on a live-fire range. (Photo credit: U.S. Army)...

17/02/2026

The AI Wild West, and Why it Needs a Sheriff

AI Is Scaling Faster Than Governance And That's a Risk AI adoption hasn't rolled out through neat transformation programmes. It has spread organically...

17/02/2026

TV Viewing Hits 12-Month High in Nielsen's January Report of The Gauge

Cable Surges 9% on Strength of College Football Playoffs and News, while NFL Delivers Top 15 Broadcast Telecasts ESPN (+86%) and FOX News Channel (+17%) Combin...

17/02/2026

Gray Signs Deal for Atlanta Braves Spring Training Games

Share Copy link Facebook X Linkedin Bluesky Email...

17/02/2026

January TV Viewing Hits 12-Month High

Share Copy link Facebook X Linkedin Bluesky Email...

17/02/2026

TAG Video Systems Expands MCS 1.7.0 With New Lens Visualization Tool

Share Copy link Facebook X Linkedin Bluesky Email...

17/02/2026

BCNEXXT Positions Broadcasters to Unlock New Revenue Oppo...

As the FCC explores a potential clawback of Upper C-Band spectrum currently used for broadcast distribution, BCNEXXT is working with broadcasters and operators ...

17/02/2026

VEON and Hala to Explore Partnership in Ride-Hailing Services

17 Feb 2026 VEON and Hala to Explore Partnership in Ride-Hailing Services Dubai, February 17, 2026 - VEON Ltd. (Nasdaq: VEON), a global digital operator ( VEON...

17/02/2026

Qube Cinema announces the acquisition of the business of Arts Alliance Media (AAM)

Qube Cinema announces the acquisition of the business of Arts Alliance Media (AA...

17/02/2026

Unveiling What's Next on Netflix for Australia and New Zealand in 2026

Back to All News Unveiling What's Next on Netflix for Australia and New Zealand in 2026 Entertainment 17 February 2026 GlobalAustraliaNew Zealand Link ...

17/02/2026

WBD Files Definitive Proxy Statement and Schedules Special Meeting for March 20, 2026, to Approve the WBD-Netflix Transaction

Back to All News WBD Files Definitive Proxy Statement and Schedules Special Mee...

17/02/2026

Arvato Systems Once Again Named One of Germany's Best IT Service Providers.

Arvato Systems Once Again Named One of Germany's Best IT Service Providers. The best IT service providers in 2026 G tersloh - Arvato Systems is once ag...

17/02/2026

ABC Consumer Media Report: Trusted data for a changing market

Open the report Our 2025 report brings together results from 155 consumer magazine titles across 42 market sectors, published by 59 media owners. Over the year...

17/02/2026

RT Choice Music Prize Announcement Classic Irish Album

Celebrating 21 Years of the RT Choice Music Prize RT Choice Music Prize In association with IMRO and IRMA Classic Irish Album And the winning album is T...

17/02/2026

RT to provide Irish-language commentary on major rugby and soccer fixtures throughout 2026

RT has today announced that as part of its public-service remit, it will expand...

16/02/2026

Live From NBA All-Star 2026: NEP Group Powers Game Capture With TFC Infrastructure, Specialty Cameras

The production infrastructure scaled seamlessly from regular-season games to the...

16/02/2026

Live From NBA All-Star 2026: NBC Sports' Ryan Soucy and Kim Titone on Ops for a New Era of NBA on NBC'

The operations team reimagined traditional workflows were reimagined and built a...

16/02/2026

Live From NBA All-Star 2026: NBA TV's New League-Run Era Debuts With Live Studio Shows

Once again an in-house operation, the league's network is producing shows fr...

16/02/2026

Live From NBA All-Star 2026: For NBA Broadcast Operations, There's a Lot of New'

The league is producing its first All-Star with NBC, Intuit Dome, and, once agai...

16/02/2026

Boland's Groundbreaking 12G 4K Video Wall

When a customer in London needed a 4K video wall that could take 12G input, the media company found no video wall suppliers that could meet their specs and time...

16/02/2026

NATO Upgrades Brussels HQ Broadcast Studio with Grass Valley Cameras

Grass Valley, a media and entertainment technology innovator, has won a competitive NATO-wide tender to provide the new camera system for NATO's main broadc...

16/02/2026

Comcast Business Powers February's Biggest Sports Moments

Comcast Business announces it is again partnering with NBCUniversal to architect and manage critical components of the linear and digital broadcast for three of...

16/02/2026

NBC Olympics on NEP's Critical Role for the 2026 Winter Olympics

Week two of the 2026 Milano Cortina Winter Olympics are underway and crucial to the NBC Sports efforts has been the NEP Group which is providing a full range of...

16/02/2026

Train Dreams, Sorry, Baby, Lurker, and Other Sundance Institute-Supported Wins at the 2026 Spirit Awards

The Film Independent Spirit Awards have officially switched things up. This Sund...

16/02/2026

SBS's Powerhouse Current Affairs Line-up is Back with the Conversations and Stories That Matter to You

SBS's Powerhouse Current Affairs Line-up is Back with the Conversations and ...

16/02/2026

L3Harris Receives New Contract to Power THAAD Interceptors

A THAAD interceptor is launched during a successful Missile Defense Agency intercept test (Photo Credit: Missile Defense Agency)...

16/02/2026

aconnic ramping up production for ACCEED 4104 10 Gigabit system for international markets

aconnic AG (ISIN: DE000A0LBKW6), Munich, is increasing capacity and ramping up p...

16/02/2026

Content Vault Expands Device Specific File Encryption to...

Content Vault, the patent-pending content security platform originally developed for the film, television and entertainment industries, has today announced a ma...

16/02/2026

Unscripted Content Producer MEGUMI Signs Exclusive Partnership with Netflix

Back to All News Unscripted Content Producer MEGUMI Signs Exclusive Partnership with Netflix Entertainment 16 February 2026 GlobalJapan Link copied to clip...

16/02/2026

Mission: Impossible: Capturing sound 10 meters underwater

In Ep62 of the AMPS Podcast, they share how a 3 mm DPA 6061 Subminiature Mic became part of a custom solution hidden inside the underwater mask capturing breath...

16/02/2026

Netflix Releases the Second Season of 'Gangs of Galicia' on April 3

Back to All News Netflix Releases the Second Season of Gangs of Galicia on April 3 Entertainment 16 February 2026 GlobalSpain Link copied to clipboard Wat...

16/02/2026

"Unscripted" Producer MEGUMI Signs Exclusive Partnership with Netflix

Back to All News "Unscripted" Producer MEGUMI Signs Exclusive Partnership with Netflix Entertainment 16 February 2026 GlobalJapan Link copied to clipboard ...

16/02/2026

RT announces the Appointments of new Clarity Correspondent and Policy & Analysis Correspondent

RT News & Current Affairs has today announced the new appointments of journalis...

16/02/2026

New Data Shows NVIDIA Blackwell Ultra Delivers up to 50x Better Performance and 35x Lower Costs for Agentic AI

The NVIDIA Blackwell platform has been widely adopted by leading inference provi...

15/02/2026

Live From NBA All-Star 2026: Entertainment Takes the Court in a Big Way

With new partnership between the league and NBC, workflows distinguish more between live, broadcast sound There'll be a lot new for the 75th NBA All-Star W...

15/02/2026

Live From NBA All-Star 2026: NBC Sports Director Pierre Moossa Previews NBC's Return to the Event

After 24-year absence, NBC Sports returns to NBA All-Star Weekend with unique ca...

15/02/2026

Live From NBA All-Star 2026: Peacock, NBC Sports Offer Viewers a Front-Row Seat With Courtside Live'

New to NBA coverage, the viewer experience offers several angles in addition to ...

15/02/2026

Live From NBA All-Star 2026: NBC Sports Returns With Plenty of Tech Toys in Tow

Coverage features 4X-slo-mo Supracam and Steadicam, Nucleus 4K cameras, closer play-by-play angle, 10 player mics NBC Sports is in the midst of its first NBA A...

14/02/2026

Cineverse Acquires TV Monetization Platform IndiCue

Share Copy link Facebook X Linkedin Bluesky Email...