01 How generative visual models actually work, told as a story 02 Orientation: the one-paragraph version, and the one-page version 03 What is a generative visual model, really? 04 The deep history, part one: from 2015 to Stable Diffusion 05 The deep history, part two: the transformer takes over 06 How diffusion really works, in plain English 07 Refinements on the basic recipe 08 From images to video: the temporal dimension 09 How audio generation fits, in brief 10 Where the training data actually comes from 11 The deep history of training datasets 12 Modern engineering: distillation, MoE, and the speed-quality frontier 13 The economics of generative imagery 14 The single most important piece of infrastructure: ComfyUI 15 Questions readers always ask 16 How research actually works 17 The landmark papers, and why they mattered 18 When generation learned to reason 19 When video learned to talk 20 Pulling it all together: the modern stack
21 Black Forest Labs (FLUX) 22 Google DeepMind (Imagen, Nano Banana) 23 OpenAI (DALL-E, GPT-Image, and GPT-Image-2) 24 Ideogram 25 Recraft 26 Midjourney 27 Stability AI (Stable Diffusion family) 28 Alibaba (Qwen-Image, Tongyi Wanxiang) 29 ByteDance (Seedream) 30 Krea (Krea-1) 31 Other notable image models 32 The mid-2026 shift: Meta arrives, the open-weights wave, and Google retires Imagen
33 What a LoRA actually is 34 The two knobs that matter: rank and alpha 35 The training dataset is where LoRAs are made or broken 36 The four main types of LoRAs and what each one is for 37 LoRAs from the artist's perspective: a working playbook 38 LoRAs from the operator's perspective: build, host, or ignore 39 The strategic implications: why LoRAs reshape the field
40 OpenAI (Sora, Sora 2) 41 Google DeepMind (Veo) 42 Kuaishou (Kling) 43 Runway (Gen-3, Gen-4, Aleph, Act-Two) 44 MiniMax (Hailuo) 45 Alibaba (Wan, Tongyi Wanxiang) 46 ByteDance (Seedance) 47 Luma Labs (Dream Machine, Ray) 48 Pika Labs 49 XAI (Grok Imagine) 50 Other notable video models 51 The mid-2026 shift: OpenAI exits consumer video, Google's Omni rebrand, and the Chinese surge
52 How audio generation actually works 53 ElevenLabs 54 The voice landscape beyond ElevenLabs 55 Suno 56 Udio 57 The music and sound landscape 58 Native audio in video, and the pressure it creates 59 The reckoning: licensing, likeness, and consent
60 Image and video editing models 61 Lipsync and audio-visual synchronization 62 3D generation models 63 Upscaling models
64 The image model landscape, in practice 65 The video model landscape, in practice 66 The audio model landscape, in practice 67 The actual decision framework 68 What benchmarks actually measure (and what they don't) 69 The numbers, what the leaderboards actually say
70 What you are actually running 71 Reading the graph 72 The default workflow, node by node 73 The recipes, part one: img2img, inpainting, outpainting, upscaling 74 The recipes, part two: LoRA, ControlNet, IPAdapter 75 Video and audio in the graph 76 The ecosystem: the Manager, custom nodes, and the registry 77 Headless and agentic: driving the graph by machine 78 Practical: memory, speed, and reproducibility 79 Local or rented: where ComfyUI runs 80 The pitfalls that cost you a weekend 81 The 80/20 of getting good 82 How the pros actually use it 83 Becoming senior
84 The dimensions of quality 85 What 'quality' means for specific use cases 86 Brand-safe generation for commercial work 87 Evaluating audio quality
88 The model is not the product anymore 89 The whole board 90 You do not use a model, you use a workflow 91 The production pipeline, end to end 92 The audio pillar 93 Distribution: where AI-native content lives and pays 94 Realtime and interactive, a different mode
95 Development: everything before a frame exists 96 Production, part one: the engines and the canvas 97 Production, part two: voice, sound, dimension, and world 98 Post-production: finishing, and the rights problem 99 Inference: the layer nobody sees 100 Distribution: where it meets an audience 101 Inside fal: one platform up close
102 What separates good from great 103 The professional generation workflow 104 Prompting at a senior level 105 Taste as the final differentiator 106 The workflows opening up right now 107 Directing audio: voice, music, and sound
108 Vibe directing: the new posture 109 The platforms and the agentic production loop 110 How an AI film actually gets made 111 The craft: consistency and control 112 The craft: technique, volume, and the edit 113 Sound and post: the underrated stage 114 The films and the filmmakers 115 Where AI wins first, by format 116 The hybrid pipeline, AI inside a live-action production 117 The studios and the money 118 Festivals, labor, and the law 119 The team of one, and the author's chair 120 The frontier, and the honest limits
121 The directors choose sides 122 The guilds and the rules 123 The deals, and what actually happened 124 The scene, the memes, and the new auteurs 125 Slop, consent, and the fight over the human
126 What the industry data actually says 127 The generative imagery stack and where you sit 128 The build vs buy decisions that define your cost structure 129 Unit economics: where the money actually goes 130 Strategic positioning: where the moats actually live 131 Legal and operational realities 132 A decision framework for your specific situation 133 Excellence in fashion ecom and film/TV, today and in the months ahead 134 The reliability gap: why professional use still hurts 135 Provenance, watermarking, and the August 2026 deadline
136 Distillation and the race to one-step generation 137 Character consistency and identity preservation 138 Long-form video and the temporal coherence problem 139 The end of the image-versus-video distinction 140 World models and interactive generation 141 World models leave the lab 142 The audio frontier: speech, music, and sound catch up 143 Where the open versus closed split goes 144 The visual roadmap 145 The state of the art, mid-2026 146 The pace of innovation and what to expect next
147 Image specialty models 148 Video specialty models 149 Audio and music
150 Further reading and primary sources
151 Glossary of Terms