How generative visual models actually work, told as a story
Chapter 01
1 min read
Reviewed v78 · August 2026
How generative visual models actually work, told as a story
Before we walk through individual models, we are going to slow down and build the entire conceptual foundation from scratch. This part of the document assumes you know nothing about machine learning, nothing about diffusion, nothing about transformers, and nothing about the people who built any of this. By the end of it, you should understand not just what these systems do, but why they do it the way they do, who figured it out, when, and what trade-offs they were navigating. Once you have that, the rest of the document, the company-by-company breakdowns, the version histories, the comparisons, stops being an exercise in memorizationMemorizationWhen a model can reproduce specific training examples on demand rather than only learning general patterns, a central concern in the copyright debate. and starts being a series of obvious inferences.
The story has a shape. It starts with a forgotten 2015 paper by a Stanford physicist that nobody read. It runs through a five-person research lab in Munich that gave the world the model that made AI image generation a household phenomenon. It picks up speed through a Berkeley graduate student whose paper was rejected from a major conference for being unoriginal, and then became the architectural backbone of Sora. It involves a single anonymous developer who built the most important piece of infrastructure in the entire ecosystem in his spare time, then turned it into a venture-backed company. And it ends, for now, with a handful of well-funded labs in San Francisco, Beijing, Hangzhou, and Freiburg, Germany, racing to figure out what comes after diffusion. We will visit all of them.
A note on jargon. Every time we introduce a new technical term, we will define it the first time it appears, in plain language, with an analogy whenever one is honest and useful. Then we will use the term normally afterward. The goal is for you to finish this document feeling like the words are familiar, not magical.
Check your understanding
pass: 5 of 7
Answer at least 5 of 7 correctly to unlock the next chapter.
1. The story of generative visual models is said to begin with what?
2. What role did the five-person research lab in Munich play?
3. What happened to the Berkeley graduate student's paper before it became the backbone of Sora?
4. How is the anonymous developer's contribution to the ecosystem described?
5. Which set of locations is named as where the story ends, with labs racing to find what comes after diffusion?
6. What is the stated payoff of building the conceptual foundation from scratch first?
7. How does the text say it will handle technical jargon?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.