
What makes a robot gripper useful isn't that it can pick up one object - it's that it can pick up the next one, and the one after that, with a tool it's never held before.
What makes an autonomous vehicle system safe isn't just that it can reason through a situation - it's that it can do so quickly enough on the hardware actually installed in the car.
What makes a virtual agent capable is exposure to as many different environments as possible before it faces the real world.
At this year's Computer Vision and Pattern Recognition (CVPR) conference, NVIDIA Research is presenting three papers that address each of these challenges - and share a common theme: training at scale creates systems that generalize across diverse applications.
The three papers cover different challenges in physical AI research:
GraspGen-X, the first foundation model for zero-shot grasping, was trained on billions of simulated grasps to work with any gripper it's shown.
LCDrive introduces a model that replaces expensive text-based reasoning with compact latent representations, letting autonomous vehicles think faster on embedded hardware.
NitroGen is a generalized gameplay AI foundation model that harnesses the NVIDIA Isaac GR00T robot foundation model architecture to help train embodied agents in virtual environments across tens of thousands of hours of interaction.
NVIDIA also unveiled at CVPR new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots and vision AI systems.
NitroGen and another NVIDIA-authored paper, PixelDIT, were named best paper finalists at the conference - an accolade given to just 15 of over 4,000 accepted papers at CVPR.
The First Foundation Model for Grasping Most AI systems for robotic grasping are specialists.
A vision-language-action policy trained for a two-finger gripper only learns to grasp with those two fingers. Similarly, a policy for dextrous grasping will only work for the bespoke multi-fingered gripper it's trained on. For every new embodiment, the process typically needs to be repeated - requiring new training data, fine-tuning and validation. This constraint means most robotics companies pick a gripper, train for it and stick with it.
GraspGen-X is the first foundation model for grasping built to eliminate this bottleneck.
Like a large language model that can apply its understanding of language to a new task without retraining, GraspGen-X applies its understanding of geometry and contact to any robotic gripper it encounters. Given the geometry of a new gripper and an unknown object it's never seen before, the model generates reliable grasp pose proposals to enable the robot to grasp the object.
https://blogs.nvidia.com/wp-content/uploads/2026/06/GraspGenX.mp4
To get there, the researchers needed a dataset that's impossible to collect in the real world at scale. They generated 2 billion simulated grasps across thousands of object shapes and synthetic gripper configurations, spanning the diversity of form factors a deployed robot might encounter.
For robot developers, this foundation model eliminates the need for per-gripper training cycles and can be applied out of the box for several commonly used grippers. GraspGenX can be used in conjunction with curoboV2, a new CUDA-accelerated motion planning library, to achieve these grasp poses in unknown environments.
Building on the GraspGen research foundation, another paper, Grasp-MPC - presented at ICRA 2026 - advances the next step in the pipeline: moving from grasp generation to closed-loop grasp execution.
Teaching Autonomous Vehicles to Think Faster In recent years, researchers have found that letting an AI reason - generating intermediate thinking steps before committing to an answer - reliably improves its decision-making.
For autonomous vehicles, the challenge is doing that reasoning on the hardware inside an actual vehicle. Text-based chain-of-thought reasoning generates words, and every word is a token that takes time to produce. On the processor running inside a car, token count is a real constraint on how fast the system can respond.
LCDrive tackles this problem by replacing words with compressed latent representations.
Instead of generating human-readable reasoning steps, the system thinks in a compact latent space - states that capture spatial information rather than producing text. The architecture alternates between two kinds of thinking: proposing candidate actions, then predicting what the world will look like if those actions are taken.
It uses that predicted world state to refine its next step. It's the same reasoning loop - just in a more computationally efficient form than natural language.
The result: comparable output trajectory quality to text-based reasoning, using roughly half the tokens.
The model was built on NVIDIA Alpamayo and trained using supervision derived from existing vehicle data.
Embodied Agents Trained in Virtual Worlds Isaac GR00T - NVIDIA's open foundation model for humanoid robots - is built on a simple principle: expose a model to enough diverse situations, and it will generalize to ones it hasn't seen.
NitroGen extends that principle to virtual environments, using the GR00T architecture to train a foundation model for embodied agents across a breadth of virtual worlds.
Video games offer something that's hard to build from scratch: structured, varied worlds with defined goals and well-specified success conditions. They're high-quality training environments, available at scale.
NitroGen treats them that way - as a training ground for agents that will eventually be trained to handle novel real- or simulated-world situations, like powering a robot that helps with housework based on broad instructions such as, Put these items away in the
North America Stories
06/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/06/2026
Matrox Video today announced the launch of the Matrox Maevex MGX Series, a new lineup of IPMX-ready video encoders and decoders with USB support that is enginee...
06/06/2026
Atomos (booth# S1006) announced immediate availability of Sumo PRO-19, a new 19-inch 4K HDR monitor-recorder-switcher designed as a central production hub for m...
06/06/2026
Cine Gear Expo 2026 DoPchoice introduces a complete light-shaping system for the new Creamsource Vortex2, further expanding the fixture's wide versatility...
06/06/2026
SmallHD will debut a new OLED 4K production monitor with a 16 display, the OLED 16, that combines OLED's industry-benchmark contrast ratio, exceptional col...
06/06/2026
Ed Lachman ASC, Caleb Deschanel ASC, and M. David Mullen ASC Being Honored at the 30th Annual Cine Gear Expo LA
Cine Gear Expo has announced the recipients of ...
06/06/2026
World premiere at Cine Gear Expo Los Angeles, June 5-6...
06/06/2026
See it first at Cine Gear Expo LA, Booth #S1703
DoPchoice introduces the latest addition to its inflatable AIRGLOW series the 8 2 Frame for RuPixel Canvas. ...
05/06/2026
The university-wide initiative has pushed their creative content to a new level
Although collegiate athletics at a single institution can contain numerous spo...
05/06/2026
In-venue and creative video staffers at the professional and collegiate level ha...
05/06/2026
Synamedia has announced that Lumine Group has agreed to acquire its Video Network business. The company is positioning the transition as the start of a new phas...
05/06/2026
Ateme, Broadcasting Center Europe (BCE), and Scaleway have announced a strategic partnership to deliver a cloud-based media supply chain covering ingest through...
05/06/2026
Audinate has announced three new additions to its Dante AVIO Install adapter series: a 4-Channel Analog Input, a 4-Channel Analog Output, and a 2-Ch In/2-Ch Out...
05/06/2026
Sportradar Group AG has announced a multi-year extension of its exclusive global...
05/06/2026
The M6 Group will broadcast FIFA World Cup 2026 matches live and in Ultra High D...
05/06/2026
The Athletic has announced that PGA TOUR highlights will be integrated into its golf coverage beginning with the Memorial Tournament presented by Workday. PGA T...
05/06/2026
When the FIFA World Cup arrives in North America in 2026, it will bring more tha...
05/06/2026
FIFA and DAZN have announced the launch of FIFA exclusively on DAZN, consolidating FIFA's content portfolio within DAZN's sports platform. The move fol...
05/06/2026
Telemundo's exclusive Spanish-language coverage of the FIFA World Cup 2026 G...
05/06/2026
Dolby Laboratories and NBCUniversal have announced that Peacock will stream Tele...
05/06/2026
Formula 1 has announced a 10-year extension to keep the Las Vegas Grand Prix on the F1 calendar through 2037. Las Vegas Grand Prix, Inc., Clark County, and the ...
05/06/2026
The broadcast-engineering team overcomes wind, speed, and salt water - and dista...
05/06/2026
Deploying both onsite and remote crews, the company is providing calibrated-came...
05/06/2026
With Inter&Co Stadium unavailable, ESPN's UFL team rebuilt its broadcast pla...
05/06/2026
One of the most exciting and informative events on the SVG annual event calendar is the Regional Sports Production Summit, an annual gathering of industry profe...
05/06/2026
For the race's third year at the historic racetrack, the broadcaster has added cameras and will incorporate multiple drones
The 158th edition of the Belmon...
05/06/2026
Ratings Roundup is a rundown of recent rating news and is derived from press rel...
05/06/2026
MRI-Simmons and S&P Global Mobility are expanding advanced audience capabilities...
05/06/2026
New Nielsen data shows insurance ad spend grew 11%, while consumers remain highl...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
ASG Promotes Joe Marchitto to Western Regional CTO
Brie Clayton June 5, 2026
0 Comments
Appointment to Support Engineering Alignment and Client Experi...
05/06/2026
Stargate Studios Colombia Uses DaVinci Resolve Studio for Vertical Microdramas
Brie Clayton June 5, 2026
0 Comments
End to end post in one platform al...
05/06/2026
People Need to Come First When We Use AI
Andy Marken June 5, 2026
0 Comments
It's just surviving. Life's very existence requires destruction....
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
05/06/2026
Frequency, the engine powering many of the world's leading streaming television channels, today announced the launch of In-Scene Advertising, a new monetiza...
05/06/2026
Berklee Study Reveals Video Has Become Essential to Music Careers Survey findings show social platforms have become the primary source of music for video cont...
04/06/2026
Sony Electronics is introducing the SRG-AS10, a 4K 60p-compatible PTZ auto-frami...
04/06/2026
This recent grad from Spring, TX, led creative-video output for the Aggies' men's basketball team last season and has been producing video and creating ...
04/06/2026
For the first time at a women's golf major, every player in the field will r...
04/06/2026
Three Panasonic PT-RQ45 40,000-lumen 3-Chip DLP projectors made their first live...
04/06/2026
Bitmovin and Akamai have announced a collaboration with NRJ Group, a French mult...