
Think of a professional athlete. What separates elite performers is what happens between games: continuous refinement, adjusting to new opponents and sharpening skills based on what the last game exposed.
Agentic AI works the same way. A model is no longer asked for an answer. It's given a goal and has to keep adapting as environments shift, edge cases emerge and tools change. Unlike a generative model responding to a prompt, an agentic model must plan, use different tools and recover from problems it encounters mid-run.
That's why post-training, the phase that refines a model after initial training on raw data, is no longer a one-time finishing step. It's continuous, because the environment that agentic models operate in shifts fast. The tools an agent uses can change week to week. Edge cases surface in production that no test set anticipated. Each deployment brings its own codebase, policies and environment.
Post-training runs loop back from production as new problems surface. The compute footprint grows not because any single run is larger, but because the runs never stop. Agentic AI introduces a new compute pattern for post-training, making it the central workload of the agentic era and the primary driver of intelligence per dollar.
The goal of post-training is to maximize intelligence per dollar by maximizing the yield of every forward and backward pass in the continuous learning cycle. The forward pass - inference - is measured in cost per token. That means that every improvement to cost per token flows directly into intelligence per dollar.
Agentic Post-Training Demystified Post-training is where intelligence is built. In pretraining, the model learns to predict the next token, which gives it fluency but not intelligence. Post-training is where it learns to write code, plan a multistep task, use a search tool and recover when something goes wrong. Inference is what comes after: the model working on the job, priced in cost per token.
Because there's no answer key to memorize, only a reward, the model learns by reinforcement learning (RL) techniques. When given a task, it writes out an attempt - the forward pass - the same work it does on the job. The attempt is scored, and the lesson updates the model's weights - the backward pass. Across millions of attempts, intelligence grows.
Each step is compute intensive, and running this loop at scale is an orchestration problem: thousands of environments generating rollouts in parallel, rewards being verified and updated weights flowing back into training with accelerators fully utilized. NVIDIA NeMo open libraries, such as NeMo Gym for training environments and NeMo RL for distributed post-training, turn post-training from bespoke research code into repeatable infrastructure.
Why Intelligence per Dollar Extends Cost per Token If inference is the revenue engine, post-training is the multiplier: the more capable the model, the higher the value of every token served.
Cost per token is the key metric for the inference factory: the all-in cost of delivering 1 million tokens. Intelligence per dollar sits one layer up, answering a different question: what does it cost to build a model worth serving, and keep it worth serving as its environment changes?
The two are nested, not competing. AI infrastructure that lowers cost per token also lowers the cost of every point of intelligence built into the model. And every point of intelligence built in raises the value of every token the inference factory serves.
In other words, cost per token measures operating yield; intelligence per dollar measures whether the investment in model intelligence is paying off.
Maximizing Intelligence per Dollar: Post-Training Nemotron 3 Ultra NVIDIA Nemotron 3 Ultra - an open weight, 550-billion-parameter mixture-of-experts (MoE) model, offers verifiable benchmarks and a fully disclosed post-training recipe run on NeMo RL. It scored 71.7% on a standard real-world coding benchmark, SWE-bench verified, where it produced a working fix for roughly seven in 10 real software bugs from open source projects, each one checked against the project's own tests.
Illustrative 20 billion rollout tokens, based on prior-generation Nemotron 3 Super's 1.2 million rollouts at 10,000 tokens each, scaled up for the larger Ultra model. Intelligence per dollar between platforms is independent of this assumption; the absolute values scale with the token count. The NVIDIA Blackwell platform lowers cost per run and makes the frequent post-training the agentic era demands economically viable. That intelligence is reaped across every token served.
The NVIDIA Vera Rubin platform extends the trajectory further, training the largest models with one-fourth the GPUs of the Blackwell generation. It was codesigned from end to end to maximize intelligence per dollar for the agentic post-training load: more rollouts per run, more environments in play and post-training cycles that never stop.
Post-Training Workflows in Action Prime Intellect's Lab continuously post-trains frontier open models on NVIDIA Blackwell and uses NVIDIA Dynamo for inference orchestration. With Vera Rubin, Prime Intellect plans to scale reinforcement learning environments, generate more rollouts per run and accelerate training-to-inference iteration loops to maximize intelligence per dollar for businesses.
Prime Intellect has optimized its sandbox infrastructure to integrate with NVIDIA Vera CPUs, enabling low-latency, energy-efficient reinforcement learning. Open source tools and models such as NVIDIA Nemotron and NVIDIA NeMo Gym are also integrated into its software stack. When comparing realistic RL sandbox workloads against alternative x86 architectures, Prime Intellect found that Vera delivers, on average, 30% greater throughput per CPU.
Perplexity's RL post-training st
North America Stories
07/08/2026
Berklee City Music Presents Nine Scholarships and Awards at Annual Concert Student recipients from seven states and Canada received scholarships, housing awar...
07/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
07/08/2026
Riedel Communications at IBC2026: Live Production Experience Showcasing Integrated IP Production Workflows
At IBC2026, Riedel Communications will bring modern ...
07/08/2026
Shotoku Showcases Expanding Aura PTZ Ecosystem and Award-Winning Swoop Robotic Crane at IBC2026
Shotoku Broadcast Systems will officially introduce its Aura fa...
07/08/2026
Highlights will include new Dante fiber connectivity, expanded IPMX capabilities and award-winning signal processing solutions
Cobalt Digital, the leading des...
07/08/2026
Bob Moses Concert at Red Rocks Captured with URSA Cine 12K LF
Brie Clayton August 6, 2026
0 Comments
PYXIS 12K and URSA Cine 12K LF enable cinematic c...
07/08/2026
Use It or Else Is Not an AI Implementation Strategy
Andy Marken August 6, 2026
0 Comments
Source - Hitchhikers Guide to the Universe, Touchstone Pi...
06/08/2026
The broadcaster's 14-race schedule begins in Iowa and ends in November with ...
06/08/2026
In-venue and creative video staffers at the professional and collegiate level ha...
06/08/2026
As media organizations consolidate teams, technology platforms, and content libr...
06/08/2026
Firmware updates introduce Dolby's next-generation picture engine, content-aware optimization, and new motion and ambient-light tools...
06/08/2026
TNT Sports will have wall-to-wall coverage of the upcoming FIBA Women's Bask...
06/08/2026
swXtch.io (Stand 5.MR13) has unveiled groundSwXtch for Audio over IP (AoIP), a software application that lets broadcasters and media organizations move audio in...
06/08/2026
At IBC2026, Riedel Communications (Stands: 10.A24, 10.A31, 10.A38) will bring modern media production to life with the Live Production Experience - a connected ...
06/08/2026
The first live-sports director to receive the prestigious honor, eight-time Emmy...
06/08/2026
NBC Sports will continue to present the Preakness Stakes on NBC and Peacock thro...
06/08/2026
Cobalt Digital (Stand 8.F90) continues to reinforce its commitment to delivering...
06/08/2026
Wave Central has promoted promotion Jeff Daubert to Director of Sales, Americas.
In his expanded role, Daubert will lead Wave Central's sales strategy and ...
06/08/2026
KMH Integration is expanding its ability to serve media, broadcast, and IT organizations by adding Keith Hanadel, an accomplished architect with decades of indu...
06/08/2026
Advanced Systems Group (ASG) has promoted Michele Ferreira from Vice President of its Systems Integration team to the newly created position of Chief Business O...
06/08/2026
Demonstrations will include new JPEG XS, Dolby audio, cloud-gateway, stream-processing, and distribution capabilities...
06/08/2026
The deal includes three premium mobile units, engineering personnel, existing br...
06/08/2026
Broadcast Management Group produced Amazon Prime's inaugural Obsessed Fest a...
06/08/2026
Exhibit will feature new variable-ND functionality, OLED viewfinders, remote-camera tools, and UHD production monitors...
06/08/2026
Compact full-frame E-mount lens targets sports, wildlife, and other long-range shooting applications...
06/08/2026
Live broadcast will honor Roger Federer and Mary Carillo on Aug. 29, with onsite...
06/08/2026
End-zone board receives higher-resolution LED technology and an updated Show Control system ahead of the 2026 football season...
06/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/08/2026
Proven offering builds on more than 20,000 live sporting events with greater automation, flexibility, and operational control
LTN , the industry leader in tran...
06/08/2026
Hornets Tech, a specialist in broadcast connectivity and IP video, is bringing a new range of encoders, decoders and an IPTV platform to IBC2026.
Broadcast in...
06/08/2026
Introducing Adobe for ChatGPT: Create, edit and get work done - all in ChatGPT
Deepti Pradeep August 6, 2026
0 Comments
The Adobe plugin in ChatGPT br...
06/08/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
06/08/2026
At IBC2026, MediaKind will make its first major appearance as a unified global powerhouse in video, showcasing one of the world's most comprehensive video i...
06/08/2026
Big Blue Marble (#5.A63) will demonstrate how its integrated technology and operational expertise help media companies scale premium services with less complexi...
06/08/2026
LA JOLLA, CA-Scripps Research has received more than $500,000 in first-year fund...
06/08/2026
LA JOLLA, CA-Viruses are masters at invading our cells thanks to specialized proteins that coat their surfaces. When scientists design vaccines, they often crea...
06/08/2026
LA JOLLA, CA-The brain has its own immune system, which detects threats and mounts a defense. A growing body of evidence has shown that in Alzheimer's disea...
06/08/2026
LA JOLLA, CA-Scripps Research chemist Jin-Quan Yu has been elected to the National Academy of Sciences (NAS), one of the highest honors a scientist can achieve....
06/08/2026
LA JOLLA, CA-Scripps Research ranked third in the inaugural 2026 Cure Innovation Index recognizing the top-performing institutes and centers across the United S...
06/08/2026
Benjamin Cravatt, the Gilula Chair of Chemical Biology and a professor of chemis...
06/08/2026
LA JOLLA, CA-Dennis Burton, professor and the James & Jessie Minor Chair in Immu...
06/08/2026
LA JOLLA, CA-Inside every human cell, proteins are constantly being tagged with small chemical modifications after they're produced. Known as post-translati...
06/08/2026
LA JOLLA-Scripps Research has established the Ian Wilson Endowed Chair, a new fa...
06/08/2026
Scripps Research's Skaggs Graduate School of Chemical and Biological Science...
06/08/2026
LA JOLLA, CA-Professor Jin-Quan Yu of Scripps Research has been elected to the Fellowship of the Royal Society, the U.K.'s national academy of sciences and ...