The living updates page tracks current model versions as they ship. This section is the durable companion to that: the primary sources, papers, and reference material worth reading directly, for anyone who wants to go past this book to the ground truth.
The following is a short list of the most valuable external sources for going deeper than this document, or for tracking the state of the field after this document becomes stale. These are the sources I would trust to be honest and up-to-date rather than promotional, in rough order of usefulness to an operator or a serious practitioner.
State of Generative Media Report, fal.ai, February 2026. The single best industry-data source for the operator perspective, based on usage data from fal's inference platform which serves more than six hundred models. Published as the primary source behind the a16z commentary summarized in the industry-data chapter of this document. Available at fal.ai slash gen-media-report-volume-1.
The State of Generative Media 2026, Jennifer Li and Justine Moore, Andreessen Horowitz, February 19, 2026. The analyst companion to the fal report. Covers the multi-model routing insight, the workflow-as-unit-of-work framing, the cost-optimization data, and the open-source customization argument. Available at a16z.com slash the-state-of-generative-media-2026.
Artificial Analysis (artificialanalysis.ai). The primary independent benchmarking platform for image and video generation models. Runs blind-test evaluations across text-to-image, image-to-image, text-to-video, and image-to-video. Where new models first show up on leaderboards before anyone knows who made them, the HappyHorse-1.0 model covered in the research frontier was first identified here.
LMArena (lmarena.ai, formerly LMSYS Chatbot Arena). The mainstream community blind-test platform for both language and image models. The platform where the GPT-Image-2 tape-codename variants appeared in early 2026. The primary source for seeing which models community testers actually prefer when they do not know which one they are rating.
Black Forest Labs technical reports at bfl.ai slash blog. The primary source for FLUX architectural details. Includes the FLUX.2 report, the FLUX.2 VAE report, and occasional research notes on training techniques.
Hugging Face Hub (huggingface.co). The central repository for open-weight models in the field. Almost every open-weight image and video model mentioned in this document can be downloaded here, including FLUX.2 [dev], Stable Diffusion variants, Wan, Qwen-Image, Helios, Gemma 4, and Netflix VOID. The model cards on Hugging Face are often the most authoritative primary source for any given open-weight release.
ComfyUI documentation and workflows (comfyui.org and the GitHub repository). The best source for actually understanding how models compose into workflows, as opposed to reading abstract descriptions of them. The node-graph structure of ComfyUI is itself an educational tool, if you are trying to understand how a specific pipeline works, finding a ComfyUI workflow for it is often faster than reading a paper.
The fal.ai blog, the Replicate blog, and the Modal blog. The inference platforms publish significantly more detail about practical operational realities than the model labs do, because their customers demand it. When you need to understand cost, latency, quantization trade-offs, or deployment patterns, these are better sources than the labs' own announcements.
Smith, Justine Moore (venturetwins on X), and Pieter Levels (levelsio on X). Three of the most reliably informed and unpromotional commentators on generative imagery in public. Moore is a partner at Andreessen Horowitz who focuses on the space. Levels is an indie founder who actively stress-tests every new model release and posts honest comparisons. Both are credited in the reporting that identified the GPT-Image-2 leak in the research frontier.
That concludes the Generative Visual Models Field Guide. If you have made it this far, you now have a working understanding of how every major image and video AI model in 2026 actually works, who built it, why they built it the way they did, and how to think about choosing between them. The science is settled enough that the next time a new model launches, you should be able to read the technical announcement, identify which architectural family it is in, infer roughly what its strengths and weaknesses will be, and decide whether to try it. That is the goal of this document, and it is the foundation for everything else you will want to do with these tools.