To make the speculation concrete, here is the field as a single forward-looking visual. Five trajectories, each grounded in research already happening today, with rough timelines that I would bet on but would not stake my life on.
Fig.diagram
■Five trajectories that are already visible. None are guaranteed. All have working prototypes today.
The honest version of these predictions: the unification trajectory is the most certain, because the architectural pieces are already in place and the labs are already shipping multimodalMultimodalHandling more than one kind of data, for example text and images together, in a single model. systems. Real-time generation is mostly an optimization problem and will happen on schedule. Identity and consistency is the area where I would be least surprised by a sudden breakthrough, somebody is going to publish a paper that solves character drift in video generation, and the entire field will adopt their technique within months.
World models is the wildcard. The Sora technical report's claim that scaling video generation produces world simulators is a strong claim that, if it turns out to be more or less correct, would change the relationship between generative AI and physical simulation, robotics, and scientific computing. If it turns out to be overstated marketing, the timeline for that trajectory slides back several years and the labs pivot to other directions. We will know within the next eighteen months.
Consolidation is the trajectory that worries me most for the health of the field. The cost of training a frontier model keeps going up. The number of organizations that can afford it stays roughly constant or shrinks. The open-weightsWeightsThe learned numbers inside a model that encode everything it knows. 'Open weights' means these are downloadable. releases continue to come but they come from a smaller and smaller set of labs (essentially Black Forest Labs, Alibaba, Tencent, and ByteDance for video). If any one of those four decides to stop releasing open weightsOpen weightsModels whose trained parameters are publicly downloadable, allowing anyone to run them locally or fine-tune them. Contrasts with closed/API-only models., and the strategic logic for stopping is real, especially as competition intensifies, the open ecosystem loses a major pillar.
The other thing to watch is what OpenAI does after the Sora shutdown. Sora was supposed to be one of OpenAI's signature consumer products, and they appear to have decided that the unit economicsUnit economicsThe per-unit costs and revenue of a product, for example the cost to generate one image against the price charged for it. of running a standalone video generation service do not work for them. That decision tells you something about how brutal the economics actually are at the frontier, even a lab with effectively unlimited capital and the best model in the field at one point chose to wind down rather than keep competing. If OpenAI cannot make video generation work as a business, the bar for everyone else is higher than it looks.
Check your understanding
pass: 5 of 7
Answer at least 5 of 7 correctly to unlock the next chapter.
1. Which of the five trajectories is described as the most certain?
2. What makes world models the wildcard trajectory?
3. Why does the consolidation trajectory worry the author most?
4. What signal is the author watching for regarding FLUX.2?
5. Which four labs are named as essentially carrying open-weights video releases?
6. What does the author say the Sora shutdown reveals about frontier economics?
7. Why is identity and consistency the area the author would be least surprised to see a sudden breakthrough?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.