PhiZero — A World Model That Speaks a Physical Language Before It Renders

— by Vivax

PhiZero, new work from the Institute of Automation, Chinese Academy of Sciences (CASIA) by Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen, Tieniu Tan, Lue Fan…

Researchers at the Institute of Automation, Chinese Academy of Sciences (CASIA) — Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen, Tieniu Tan, Lue Fan and Zhaoxiang Zhang — have released PhiZero, a world model built around a "physical language": a compact, discrete representation of world-state transitions learned self-supervised from in-the-wild videos, with no labels and no simulators. Before generating a single pixel of future video, PhiZero first works out what will happen as a sequence of physical-language tokens, and only then renders that plan into frames. Reason first, render second.

The paradigm splits the labor that pixel-space predictors entangle: the reasoning stage manipulates a small, structured description of dynamics, and the rendering stage turns a settled plan into appearance. Validated on both video-generation and video-understanding benchmarks, the physical-language bottleneck helps rather than hurts — and it unlocks interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer, where motion inferred from one scene is re-rendered onto another without retraining.

For medicine, the architecture is the message: a clinical world model should reason over state transitions first, in an explicit representation clinicians can inspect, and render outputs — summaries, simulations, explanations — second. A model that can say what it expects to happen before it dresses the prediction up is a model you can audit.

Back to all news | Vivax Home

0%