The Physical AI Report Card: Physics-IQ, PAI-Bench, and Post-Training Cosmos 3 in a Day
— by Vivax
Do video models actually understand physics? Physics-IQ — a benchmark from Google DeepMind and INSAIT spanning solid mechanics, fluid dynamics, optics,…
How do you know whether a video model actually understands the world it renders? In early 2025, researchers from INSAIT and Google DeepMind proposed a blunt answer: film a real physical event, hand the model the first three seconds, and ask it to predict the rest. That is the Physics-IQ benchmark — 396 real videos, no synthetic scenes, spanning solid mechanics, fluid dynamics, optics, thermodynamics and magnetism. Because every clip has a filmed ground truth, the score is not a vibe check; it measures whether the model's imagined future matches what physics actually did. The headline finding of the original paper was sobering: visual realism and physical understanding were largely uncorrelated — models produced gorgeous, plausible-looking video that broke conservation laws within seconds.
The scoring is refreshingly mechanical. Four metrics — where action happens, when and how much, and how it unfolds — are aggregated into a single Physics-IQ score, normalized so that 100 equals the variability between two real recordings of the same event. A perfect score means the model is indistinguishable from reality re-running itself. When the paper landed, the best model of the day scored around 24.1 out of 100.
A year and a half later the leaderboard tells a story of real progress — and a long road ahead. As of mid-2026, Magi-1 leads with 56.37, ahead of VideoPoet's 29.5, with well-known systems like Sora far lower. The gap between the best 2025 score and the best 2026 score roughly doubled, driven partly by physics-aware post-training recipes. But even the leader captures barely half of real-world physical variability — no model has crossed 60 on a benchmark where 100 means parity with reality.
Physics-IQ measures one thing deeply; NVIDIA's Physical AI Bench (PAI-Bench) measures the whole stack broadly, grading world models on domain fidelity, visual quality, and physics for both generation and understanding tasks. The July 2026 generation leaderboard produced a striking result: Cosmos3-Super (83.9) and Cosmos3-Nano (83.7) now score above real source video (82.6) on the aggregate, with Veo-3 close behind. That doesn't mean the models beat reality; it means the rubric rewards their cleaner, domain-matched outputs — and it shows how fast frontier world models are converging on the benchmark's ceiling.
The most practical piece of news came from NVIDIA's developer blog: post-training Cosmos 3 no longer takes a research team a month. Using TAO agent skills, a coding agent runs LoRA fine-tuning, evaluation, and an AutoML hyperparameter sweep from a few natural-language prompts — compressing a specialist workflow into roughly a day. In their worked example on the Woven Traffic Safety dataset, the post-trained Cosmos 3 Reasoner jumped from 32% to 63.7% accuracy, learning to use cues like a distant traffic signal the base model ignored.
We read these leaderboards the way clinicians read lab panels: no single number tells the truth, but the trend does. Vivax is building clinical world models — systems that must respect the physics and physiology of real patients, not just render plausible pictures of them. Physics-IQ's lesson, that realism and understanding are different skills, is exactly why we anchor our models to measured clinical ground truth. And the Cosmos 3 post-training story matters just as much: if adapting a frontier world model to a niche domain now takes a day of agent-driven fine-tuning instead of a quarter of engineering, the advantage shifts to whoever holds the best domain data. In medicine, that is the entire game.