Contents

140 / 153

The Research Frontier

The end of the image-versus-video distinction

Chapter 139

1 min read

Reviewed v78 · August 2026

Right now, image models and video models are mostly separate. You use FLUX for images and Wan for videos. The architectures are similar but the trained models are distinct. This distinction is starting to dissolve.

The newer architectures are increasingly unified, Sora 2, Veo 3, Kling 3.0, Seedance 2.0 are all described as multimodal models that handle text, images, and video in a single network. Wan 2.5 has audio integration. Qwen-Image-Layered handles structured layered output. The trend is toward models that can take any combination of inputs (text, image, video, audio) and produce any combination of outputs, with the modalities being interchangeable rather than fixed.

By 2027, we expect most major frontier models to be 'omni-modal' in this sense, and the question 'is this an image model or a video model' will sound as quaint as 'is this a black-and-white camera or a color camera.' The capabilities will be unified and the trade-offs will be about specific feature support and quality on specific tasks.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. How are image and video models related right now, according to the chapter?

  2. 2. What is the trend in newer architectures like Sora 2, Veo 3, Kling 3.0, and Seedance 2.0?

  3. 3. What does Wan 2.5 add, per the chapter?

  4. 4. What is the described trend for inputs and outputs across modalities?

  5. 5. What does the chapter expect most major frontier models to be by 2027?

  6. 6. What analogy is used for how quaint the image-versus-video question will sound?

  7. 7. Once modalities unify, what will the trade-offs be about?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.