Contents

141 / 153

The Research Frontier

World models and interactive generation

Chapter 140

5 min read

Reviewed v78 · August 2026

The most speculative frontier is the one OpenAI's Sora technical report tried to claim, that scaling video generation models eventually produces something more than a video generator. It produces an implicit model of physics, object permanence, scene composition, and the temporal logic of the real world. Sora's report called this 'video generation models as world simulators,' and the claim has been controversial since the day it was published.

The strong version of the claim is that a sufficiently large video model has internalized enough about how the physical world works that you can use it as a simulator, for robotics training, for scientific computation, for counterfactual reasoning. The weaker version is that video models exhibit some emergent world-modeling behaviors but are nowhere near reliable enough to be used as actual simulators. The empirical state of evidence in 2026 is somewhere between these, the best video models can produce footage that respects basic physics most of the time, but they fail in characteristic ways (objects passing through each other, gravity reversing, conservation of mass being violated) often enough that calling them 'simulators' is generous.

Fig.diagram
FIXED CLIPrender once, watch it backWORLD MODEL: THE LOOPACTIONsteer / inputPREDICTnext frameFRAMEshown to youSTATEworld updatedYou are inside it, not watching it.Each action changes the next frame. Generation becomes a steerable simulator.
A world model closes the loop. Instead of rendering a fixed clip it takes an action, predicts the next frame, updates its internal state, and waits for the next action. Video generation becomes an interactive simulator you can steer.

The most interesting work in this direction is happening at Google DeepMind, at Fei-Fei Li's World Labs, and at a handful of smaller research groups. DeepMind's Genie 3 is the current frontier for interactive video environments, it takes a text prompt or image and generates a navigable space that the user can explore in real time, with the model maintaining spatial consistency as the viewer moves through it. Genie 3 is the third major iteration of this work (Genie 1 was a 2D environment generator, Genie 2 added richer 3D navigation in late 2024), and it is explicitly positioned by DeepMind as a step toward generative game engines and robotics simulation environments.

World Labs, founded by Fei-Fei Li in 2024, shipped Marble in late 2025, which is the clearest product expression of the world-model idea to date. Marble takes a single image or text prompt and generates a persistent, interactive 3D environment, you can walk around in it, interact with objects, save the state, and come back to it later with the world in the same configuration you left it. The fal/a16z State of Generative Media 2026 report explicitly flagged Marble as the moment world models moved from 'prototype to product,' meaning that for the first time a generative 3D environment was good enough to ship to users rather than just demonstrate in a paper. Marble is currently positioned for creative and game development use cases, but the technology obviously generalizes to robotics training, architectural visualization, and education as the quality improves.

The other players to watch are OpenAI (which has reframed its internal compute spend toward world simulation research as part of the stated rationale for the Sora shutdown), Meta (whose Mango model under Alexandr Wang's Superintelligence Lab is expected to push on world-model capabilities), and the Chinese labs (Tencent's HY-WorldPlay released RL post-training code in March 2026 for building real-time interactive world models on top of HunyuanVideo). This is the most contested research direction in generative imagery right now, and the labs investing in it are betting that world models are the next major category after video generation, not the same category scaled up.

If world models work, really work, in the sense that they become reliable enough to use as simulators, they change a lot. Robotics training stops requiring physical robots and uses simulated worlds instead. Game engines become obsolete for some use cases, replaced by generative world models that produce environments on demand. Architectural visualization, scientific simulation, and education all get new tools. The economic implications are enormous, but the technical question of whether the scaling story actually works is still open. The honest answer in 2026 is that we will find out within the next two years, and the answer will reshape the field one way or the other.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. What is the strong version of the world-model claim?

  2. 2. In the dollhouse analogy, what represents a world model?

  3. 3. How do today's best video models characteristically fail as simulators?

  4. 4. What does DeepMind's Genie 3 do?

  5. 5. Why did the fal/a16z report flag World Labs' Marble as significant?

  6. 6. Why are the labs working on world models mostly the same as those working on AGI?

  7. 7. If world models become commercially important, what will the field likely resemble?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.