New: the free Generative AI Weekly Field Report, everything that moved in one email a week. Read the latest issue
/ A comprehensive guide to
Generative AI
Image, Video, Audio, and the New Filmmaking
How the image, video, and audio models work, how the whole creative stack fits together, and how AI filmmaking is being made. For the working artist and the operator of a generative media company.
- Version
- 78
- Updated
- 2026.08.11
- Length
- 175K words
- Sections
- 169
- Read
- ~13 hr
PROMPTED BY Alex Park · Automata Labs
RESEARCHED & AUTHORED BY Claude · Anthropic
Recent updates · v78
All updates →- UPDMid-August refresh: frontier video goes open, and AI music makes peace2026.08.11 · v78
- UPDEarly August refresh: the law arrives, and video keeps absorbing audio2026.08.03 · v77
- UPDHybrid filmmaking: the stage, the composite techniques, and the finish2026.07.31 · v76
- UPDInline source links across the book2026.07.31 · v75
- UPDThe hybrid filmmaking process, as a working guide2026.07.31 · v74
Today's session
PrefacePreface
About this book
3min · then you're done for today
Contents
- 00Preface, about this bookStart here
Book One
The Machine
How generative models actually work, and how to wield them.
Foundations01
The Foundations
- 01How generative visual models actually work, told as a story
- 02Orientation: the one-paragraph version, and the one-page version
- 03What is a generative visual model, really?
- 04The deep history, part one: from 2015 to Stable Diffusion
- 05The deep history, part two: the transformer takes over
- 06How diffusion really works, in plain English
- 07Refinements on the basic recipe
- 08From images to video: the temporal dimension
- 09How audio generation fits, in brief
- 10Where the training data actually comes from
- 11The deep history of training datasets
- 12Modern engineering: distillation, MoE, and the speed-quality frontier
- 13The economics of generative imagery
- 14The single most important piece of infrastructure: ComfyUI
- 15Questions readers always ask
- 16How research actually works
- 17The landmark papers, and why they mattered
- 18When generation learned to reason
- 19When video learned to talk
- 20Pulling it all together: the modern stack
- The Engines
02
Image Models, by Company
- 21Black Forest Labs (FLUX)
- 22Google DeepMind (Imagen, Nano Banana)
- 23OpenAI (DALL-E, GPT-Image, and GPT-Image-2)
- 24Ideogram
- 25Recraft
- 26Midjourney
- 27Stability AI (Stable Diffusion family)
- 28Alibaba (Qwen-Image, Tongyi Wanxiang)
- 29ByteDance (Seedream)
- 30Krea (Krea-1)
- 31Other notable image models
- 32The mid-2026 shift: Meta arrives, the open-weights wave, and Google retires Imagen
03
The LoRA Deep Dive
- 33What a LoRA actually is
- 34The two knobs that matter: rank and alpha
- 35The training dataset is where LoRAs are made or broken
- 36The four main types of LoRAs and what each one is for
- 37LoRAs from the artist's perspective: a working playbook
- 38LoRAs from the operator's perspective: build, host, or ignore
- 39The strategic implications: why LoRAs reshape the field
04
Video Models, by Company
- 40OpenAI (Sora, Sora 2)
- 41Google DeepMind (Veo)
- 42Kuaishou (Kling)
- 43Runway (Gen-3, Gen-4, Aleph, Act-Two)
- 44MiniMax (Hailuo)Upd
- 45Alibaba (Wan, Tongyi Wanxiang)
- 46ByteDance (Seedance)Upd
- 47Luma Labs (Dream Machine, Ray)
- 48Pika Labs
- 49XAI (Grok Imagine)Upd
- 50Other notable video models
- 51The mid-2026 shift: OpenAI exits consumer video, Google's Omni rebrand, and the Chinese surge
06
Specialized Models: Editing, Lipsync, 3D, and Upscaling
07
Comparative Analysis: How These Models Actually Differ
- The Tooling
08
The ComfyUI Field Manual
- 70What you are actually running
- 71Reading the graph
- 72The default workflow, node by node
- 73The recipes, part one: img2img, inpainting, outpainting, upscaling
- 74The recipes, part two: LoRA, ControlNet, IPAdapter
- 75Video and audio in the graph
- 76The ecosystem: the Manager, custom nodes, and the registryUpd
- 77Headless and agentic: driving the graph by machine
- 78Practical: memory, speed, and reproducibility
- 79Local or rented: where ComfyUI runs
- 80The pitfalls that cost you a weekend
- 81The 80/20 of getting good
- 82How the pros actually use it
- 83Becoming senior
- Judgment
Book Two
The Medium
The world being built on top of the machines: the stack, the whole ecosystem, and the new filmmaking.
The Big Picture10
The Stack: Products, Workflows, and the Whole Board
11
The Ecosystem: A Field Guide to the Whole Board
- 95Development: everything before a frame exists
- 96Production, part one: the engines and the canvas
- 97Production, part two: voice, sound, dimension, and world
- 98Post-production: finishing, and the rights problem
- 99Inference: the layer nobody sees
- 100Distribution: where it meets an audience
- 101Inside fal: one platform up close
- The Craft
- The New Filmmaking
13
AI Filmmaking: Vibe Directing and Agentic Production
- 108Vibe directing: the new posture
- 109The platforms and the agentic production loop
- 110How an AI film actually gets made
- 111The craft: consistency and control
- 112The craft: technique, volume, and the edit
- 113Sound and post: the underrated stage
- 114The films and the filmmakers
- 115Where AI wins first, by format
- 116The hybrid pipeline, AI inside a live-action production
- 117The studios and the money
- 118Festivals, labor, and the law
- 119The team of one, and the author's chair
- 120The frontier, and the honest limits
- The Industry
- The Business
15
The Operator's Playbook
- 126What the industry data actually says
- 127The generative imagery stack and where you sit
- 128The build vs buy decisions that define your cost structure
- 129Unit economics: where the money actually goes
- 130Strategic positioning: where the moats actually live
- 131Legal and operational realities
- 132A decision framework for your specific situation
- 133Excellence in fashion ecom and film/TV, today and in the months ahead
- 134The reliability gap: why professional use still hurts
- 135Provenance, watermarking, and the August 2026 deadline
- The Frontier
16
The Research Frontier
- 136Distillation and the race to one-step generation
- 137Character consistency and identity preservation
- 138Long-form video and the temporal coherence problem
- 139The end of the image-versus-video distinction
- 140World models and interactive generation
- 141World models leave the lab
- 142The audio frontier: speech, music, and sound catch up
- 143Where the open versus closed split goes
- 144The visual roadmap
- 145The state of the art, mid-2026
- 146The pace of innovation and what to expect next
Back matter
Reference
Lookups, further reading, and the glossary.
Reference17
Additional Notable Models and Tools
18
Further Reading and Primary Sources
The weekly briefing
Keep up with generative AI, in one email a week.
What actually moved across image, video, and audio AI, and the research that matters, in plain English with every claim sourced. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.
Or browse past issues on the reports page.
Sourced and cited
63 tracked sourcesEvery update is checked against primary sources and ships with its citations. We prefer official announcements, papers, and independent leaderboards, and we keep a standing blocklist of the content farms that invent releases.
Image models
Video models
audio-models
Research
Industry & legal
Ecosystem
release-feeds
regulation-legal