worldmodeldata: The Data Layer for World Models, Built from AAA Games
— by Vivax
worldmodeldata is building the data layer for world models: millions of multi-modal, action-conditioned trajectories extracted from AAA video games —…
One of the quiet truths of the world-model era is that the bottleneck is no longer architecture — it is data. Enter worldmodeldata, a company building what it calls the data layer for world models. Its premise is elegant: video games are safe, controlled worlds where every interaction can be captured perfectly, so the company extracts millions of multi-modal, action-conditioned trajectories from top-performing AAA and AA titles and delivers them as training-ready datasets for labs building world models, vision-language-action systems, and physical AI. What ships is not a pile of clips but a synchronized stack: temporally aligned video, telemetry inputs, and ground-truth 3D state — every stream aligned tick to frame.
The pitch turns on a distinction anyone training a world model will recognize: generic video corpora teach a model what scenes look like, whereas engine-grade data teaches a model what a world is — a coupled system of state, motion, intent, and consequence, frame by frame, provably aligned. Three design choices carry that claim. Ground-truth 3D state comes straight from the engine, not inferred from pixels: exact per-object position, rotation, and velocity for every entity, with no perception noise, labeling error, or occlusion guesswork. Action conditioning — the hardest signal to source, because the next state depends on the action taken — is recorded, not inferred: raw controller and keystroke inputs aligned to the exact frame and state they produced. And the game state is dense: over 200 structured properties per frame, from player position and velocity vectors to view matrices, hitbox intersections, ballistic trajectories, and navmesh polygons — each one a real column in the delivered dataset.
worldmodeldata engineers for the three properties it argues decide whether a dataset can train a frontier world model: volume, diversity, and quality. On volume, today's largest public gameplay datasets contain tens of thousands of hours; the company is building toward one million hours, with petabytes of synchronized capture and billions of ticks of gameplay. On diversity, hundreds of titles, thousands of maps, varied physics, lighting conditions, and edge-case scenarios reduce the risk of overfitting to any single environment. On quality, the data is zero-loss engine output obtained under formal licensing agreements rather than scraped from the web — no hallucinated labels, no alignment drift, every frame the same source of truth the game itself runs on. Framework customers can even request additional titles, custom property extraction, or live engine instrumentation for novel research environments.
The company is also engaging the research community directly. In 2025 it published a position paper, 'Game-Generated Data, An Untapped Resource for Advanced AI Training,' examining how game-generated data could advance AI training — with a particular dive into the Joint-Embedding Predictive Architecture (JEPA), the latent-predictive family of world-model architectures championed as a path toward human-level intelligence — and it continues to pursue collaborations with academics working on world models, physical AI, and self-supervised learning. The product itself is delivered as a layered stack per hour of gameplay: 1080p multi-view video timestamped to engine ticks; full telemetry and semantic game state per tick; guaranteed temporal alignment across streams; and labeled controller, mouse, and keyboard inputs aligned to frames — recorded, not inferred.
What worldmodeldata is doing for games, Vivax is doing for medicine: our Vivax Data Layer applies the same engine-grade philosophy — state, intervention, and consequence, temporally aligned and quality-audited — to clinical streams, so that medical and clinical world models can learn from data that captures how treatment decisions actually change patient state over time. The convergence is no accident: whether the world is a game engine or a hospital ward, a world model is only as good as the aligned state-action-outcome data it learns from. We read worldmodeldata's work as strong validation that the data layer is where the next competitive edge in world models will be won — and we are glad to see the ecosystem building it. Learn more at worldmodeldata.com, or explore the Vivax Data Layer on our models page.