VERSION 1Early 2026
Initial draft. Covered every major image and video model on Krea and Flora as of its writing, organized by company. Nine parts: foundations, image models, video models, specialized models, comparative analysis, future directions, additional models, glossary, version reference. Roughly 30 pages, ~13,800 words.
VERSION 2Early 2026
File recovery and re-presentation. The V1 document was misplaced in the conversation; this version was recovered from disk and re-shared.
VERSION 3Early 2026
Major expansion to ~73 pages, ~32,500 words. Rewrote the foundations as a 10-chapter narrative with the sculptor analogy, the full historical arc from Sohl-Dickstein 2015 through the William Peebles DiT story, a dedicated ComfyUI chapter, and a dedicated economics chapter. Added founder narratives, funding histories, and strategic context to every company section in Parts 2 and 3. Defined all terminology inline. Added analogies and metaphors throughout.
VERSION 4April 11, 2026
Premium visual redesign and content expansion. Added twelve hand-authored SVG diagrams covering: the diffusion process, latent diffusion, U-Net vs Diffusion Transformer, MM-DiT dual-stream architecture, the 3D VAE for video, a basic ComfyUI workflow, a decade-long visual timeline 2015-2026, the end-to-end training pipeline, the research-to-product pipeline, the speculative future timeline 2026-2030, and a Venn diagram of the moat (compute + data + talent). New chapters added to the foundations: where training data actually comes from (the LAION origin story, the Re-LAION CSAM cleanup, the video data problem and how labs solve it), how research actually works (the five-stage pipeline from PhD desk to your screen), and the seven landmark papers with full context. New benchmarking tables in the comparative-analysis part: image leaderboard, video leaderboard, training cost comparison, lab funding and valuations. Inside-scoop callouts throughout Parts 2 and 3 with the 'real reason' behind specific company decisions. New visual roadmap section added. Premium typography refinements including pull quotes, eyebrow labels, gold accent rules, justified body text, and a version footer with page numbers. Title page redesigned with hero image, version block, and credits.
VERSION 5April 11, 2026
Reframed for dual audiences: the working creative artist who wants to become great at making images and video, and the founder, CEO, or CTO of a generative imagery company who needs to make strategic and operational decisions. Added five major new content sections built from a five-expert-perspective gap analysis (working AI artist, founder/CEO, CTO/ML infra lead, academic researcher, creative director). The new sections are: the video-models part.5 The LoRA Deep Dive (the technique from five angles, math, training, dataset prep, types, artist's playbook, operator's playbook, strategic implications, with a hand-authored architecture diagram), The Craft of Generation (ten dimensions of good vs great, the five-stage professional workflow, prompting at a senior level, taste as the final differentiator), The Operator's Playbook (the six-layer stack, the five build-vs-buy decisions, unit economics with concrete numbers, strategic positioning and moats, legal and operational realities, a decision framework), the evaluating-quality part Evaluating Quality (eight dimensions of quality, what quality means for specific commercial use cases, brand-safe generation), and the research-frontier part The Research Frontier (distillation and one-step generation, character consistency, long-form video, world models). Four new diagrams: LoRA architecture, the five-stage workflow, the operator stack, and one expanded operator economics view. New typography primitives including executive summary boxes, key takeaways boxes, ornament dividers, and drop caps. The document is now 50,000+ words and is structured as both a field guide and a practitioner's handbook.
VERSION 6April 11, 2026
Major design and structural refresh. The document was renamed from 'Generative Visual Models Field Guide' to 'A Comprehensive Guide to Image and Video AI Models.' The primary typeface was switched from Calibri to Lora, a contemporary calligraphic serif that transforms the reading experience from corporate Word document to premium publication. A Table of Contents was added at the front covering all parts, chapters, and sub-sections. The glossary of terms sits in the reference section at the back, so terminology is easy to look up.
VERSION 7April 12, 2026
Major editorial and design refresh. A full-bleed magazine-style cover replaced the V6 title page, and thirteen new part covers were added with unique thematic motifs. Margins were widened and body text enlarged for mobile reading, all 757 em-dashes were replaced, 92 acronyms expanded, and the glossary terms bolded. A new Orientation chapter was added to the top of the foundations, plus new the training-data chapter on how training data is actually collected and new the common-questions chapter on the questions readers always ask. Twenty-nine scaffolding callouts were planted at the technically densest spots across the foundations, the LoRA Deep Dive, the Operator's Playbook, the Research Frontier, and the Craft of Generation.
VERSION 8April 13, 2026
State-of-the-art refresh. A new the research frontier was added titled 'The state of the art,' a snapshot of the frontier as of this writing. It covers the Sora shutdown announced March 24, the current video SOTA hierarchy (Veo 3.1, Kling 3.0, Seedance 2.0, Wan 2.6), the FLUX.2 family, the April 4 GPT-Image-2 leak on LM Arena (later confirmed by the official April 21 launch), the open-source surge led by Qwen 3.5 Small and Alibaba's HappyHorse-1.0, and the rumored upcoming releases from Meta, OpenAI, and DeepSeek. the foundations's chapter numbering was also cleaned up so the three former 'extras' chapters (9.5, 9.7, 9.8) now read as the natural Chapters 11, 12, and 13. The document version was bumped to 8.
VERSION 9April 13, 2026
Editorial audit pass. The redundant coverage of the LAION story was merged so the training-data chapter owns the mechanical data pipeline and the training-datasets chapter owns the historical depth. A new the industry-data chapter was added at the top of the Operator's Playbook grounding the entire section in the fal.ai and a16z State of Generative Media 2026 report, with five cited findings (median 14 models per enterprise, workflow as unit of work, 58 percent cost-optimization primary, three verticals leading adoption, open source winning on customizability). The the image-models part thin sections (Midjourney, Google DeepMind, Ideogram) were rewritten to the full 5-section template, and the video-models part thin sections (Veo, Seedance) got the same treatment. A new 3.10 xAI Grok Imagine section was added since Grok was flagged as a tier-one video player by the a16z report but was missing from the document, and the research frontier World Models was updated to cover Marble from World Labs and Genie 3.
VERSION 11May 2, 2026
Major expansion across four operator-critical topics that were underserved in earlier versions. A new the visual-roadmap chapter on the pace of innovation covers empirical release cadences, a quarter-by-quarter forecast through Q1 2027, and plausibility scenarios for 2027, 2028, and 2030. A new the fashion and film chapters on excellence in fashion e-commerce and film and TV walks through vertical-specific mastery today and trajectory over the next 6 to 12 months for both, with tooling landscapes, five-to-six component frameworks, and honest forecasts. the generative-imagery-stack chapter Layers 3 and 5 were expanded with substantial new coverage of what inference platforms and application platforms are actually building beyond serving models (workflow primitives, managed fine-tuning, multi-model chaining, real-time canvas, collaborative projects, brand kit integration). A new the reliability chapter on the reliability gap covers the four professional-use failure modes (prompt adherence on do's and don'ts, product hallucination, the uncanny valley, and persistent anatomical failures) with root-cause explanations, current workarounds, and honest timelines for each. May 2026 update: GPT-Image-2 launched April 21 with record +242 Arena lead and autoregressive architecture, confirming the shift from diffusion. DALL-E 2 and 3 retiring May 12. Seedance 2.0 copyright controversy added. Grok Imagine 1.0 updated from feature to named product. Vision Banana, LTX-2, Novi AI Long Video Agent, and Adobe Creative Agent added. Q2 2026 forecast updated with confirmed events.
Table of Contents
VERSION 12July 22, 2026
Living-edition update covering the field from May 2 to July 22, 2026. the ComfyUI chapter (ComfyUI) gains a full deep dive on the KSampler: what steps do and where the returns flatten, how CFG scale trades prompt adherence against image quality, and how to choose a sampler and scheduler, with the through-line that settings must match the model class. Two new field-update sections record the mid-2026 shifts: in image models, Meta's arrival with Muse Image, the open-weights wave of Ideogram 4.0, Krea 2, and HiDream-O1, and Google's retirement of the Imagen brand; in video, OpenAI's exit from consumer video and the September 24 Sora API sunset, Google's Gemini Omni rebrand, and the surge from Kling, Seedance, and HappyHorse.
VERSION 13July 22, 2026
Comprehensiveness pass covering paradigm shifts the earlier editions underweighted. New Foundations chapters on reasoning-augmented generation (models that plan and self-check before they render) and on native audio-visual generation (synchronized sound produced in the same pass as the picture). A new Operator's Playbook section on provenance, watermarking, and the EU AI Act Article 50 transparency deadline of August 2, 2026. A new section promoting world models from research frontier to shipping product. New subsections on agent-driven ComfyUI and on Blackwell-era hardware economics and current pricing. Every claim verified against primary sources before inclusion.
VERSION 14July 22, 2026
Structural reframe integrating the ecosystem view. A new anchoring part, The Stack, sits after the foundations and retells the field from the point of view of the whole creative stack rather than individual models: why the model has become a commodity input and value has migrated to the workflow and distribution layers, a map of the five phases, the aggregator and agent layer where creators actually work, the production pipeline as a chain of specialized jobs, the audio pillar the guide had underweighted, distribution and the vertical-microdrama economy, and the realtime and interactive mode that breaks the batch-generation frame. The image and video model parts are reframed to sit inside this larger picture. Every claim was verified against primary sources first.
VERSION 15July 22, 2026
Structural overhaul and a new title. The book is regrouped into named acts that tell the 2026 story in order, from Foundations through The Big Picture, The Engines, The Craft, The Business, The Frontier, and Reference, replacing the accreted decimal part numbering. The glossary moves from the front to the reference section at the back. A new part on AI filmmaking is added to The Craft act, covering vibe directing, the agentic film platforms led by Runway and OpenArt, and the authorship debate that dominated Cannes in 2026. Every claim was verified against primary sources first.
VERSION 16July 22, 2026
A new act, The Industry, adds the beat-reporter layer the guide was missing. Five chapters cover how the field landed in Hollywood and the culture: the directors choosing sides, from Scorsese partnering with an AI lab to del Toro refusing outright; the guild and Academy rules the strikes established; the studio deals and what really happened to them, from Lionsgate-Runway to Netflix's quiet three hundred titles to the Disney-OpenAI collapse; the native creator scene, its memes, and its new auteurs; and the twin fights over slop and consent. Every quote and figure was verified against reputable reporting, with creator-claimed numbers labeled as such and unverified items excluded. Also fixes the book's numbering: legacy in-title numbers are removed and cross-references are now by name.
VERSION 22July 2026
Deepened the video-models part, which had fallen behind the image-models part in both depth and currency. Added sourced sections on Runway's 2026 turn (Aleph 2.0, the Media Router, and the Lionsgate and Bertelsmann deals), Google's move from a standalone Veo line to the Gemini Omni family, and Kuaishou's move to spin out Kling. Reconciled stale flagship and leaderboard claims across Veo, Luma, and Wan, and added dated notes on Seedance 2.5's thirty-second single take, Hailuo 2.3, Pika's slide, and the aftermath of Sora's exit. Every added claim carries at least one tier 1 to 2 source.
VERSION 23July 2026
Deepened the Ecosystem field guide, which read as a good tour but a thin one. Each phase now names more of the real players and, where it matters, carries the mid-2026 state of the board: Google's move to the Gemini Omni family, the inference layer's funding surge (fal's AWS deal at 2.5 million developers, Baseten near a 13 billion dollar valuation, Together AI at 8.3 billion), the rise of real-time world models (Reactor, World Labs' Marble, Genie 3), and the microdrama economy. Sourced where a specific claim is new.
VERSION 24July 2026
Currency pass across the image-model, leaderboard, and research-frontier sections flagged by a staleness audit. Adds dated notes bringing the standings to July 2026: GPT Image 2 at the top of the image arena, FLUX 3's multimodal launch, Nano Banana 2's tiers, Ideogram 4.0's open weights, Seedream 5.0 Pro, and, in video, Gemini Omni Flash leading both arenas. The stale spring orderings are now explicitly marked as history. Every added claim carries a tier 1 to 2 source.
VERSION 25July 2026
Language and currency cleanup: removed every self-referential 'as of April' timestamp, which had the book narrating a moment three months stale as if it were the present. Clock references now read mid-2026, genuine historical event dates read early 2026, and a few claims frozen at the April snapshot (Sora still running, Veo 3.1 on top, GPT-Image-2 still in testing, a Veo 4 expected) were corrected to what actually happened.
VERSION 26July 2026
Continued the currency audit: fixed the remaining self-referential clocks ('as of early 2026', 'as of this document's writing', 'at the time of writing') to read mid-2026, corrected a stale claim that Veo 3.1 and Kling 3.0 topped video, and folded in the mid-2026 deal and litigation news. New sourced notes cover the studio and rights-holder deals (Google/A24, Lionsgate/Runway, Runway/Bertelsmann, Getty/OpenAI), the state of the major AI copyright suits (the MiniMax case proceeding, the New York Times sanctions motion), and the EU AI Act's August 2 enforcement date.
VERSION 27July 2026
Resequenced the Foundations part into a cleaner reading arc. The conceptual build now runs orientation, history, how diffusion works, refinements, then the two data chapters together, then video, engineering, economics, and ComfyUI. The FAQ, research-process, landmark-papers, and recent-shift chapters move to a going-deeper cluster near the end, and 'Pulling it all together' is now the actual capstone that closes the part. Transitions that referenced neighbors were rewritten so the page-to-page flow still holds.
VERSION 28July 2026
Fact-verification pass plus two new figures. An independent check confirmed the frontier-snapshot claims that looked shakiest (Vision Banana, Novi AI, Helios, Gemma 4, Netflix VOID, DeepSeek V4, Spectrum) are all real and sourced, and corrected two precision errors it surfaced: Gemma 4's local memory footprint (the 12B needs roughly 16GB, not 5GB) and its date, and DeepSeek V4's Ascend claim (only inference runs on Huawei silicon; training still uses NVIDIA). Added diagrams to the distillation and Mixture-of-Experts sections, which had none.
VERSION 29July 2026
Fully integrated the ecosystem map into the book. The Ecosystem field guide now has a subsection for every product category on the board, twenty-five in all, each explaining what the category is, why it exists in the pipeline, and profiling the companies doing the most interesting work in it, from story development and storyboarding through the realtime, aggregator, canvas, and multi-shot layers around the engines, the voice, music, 3D, world-model, agentic, ads, and interactive categories, the post-production finishing stack, the inference layer, and distribution. Grounded in the ecosystem registry; honest about which names matter and which are long-tail.
VERSION 30July 2026
Broadened the Inside fal chapter to cover more of the niche, smaller-company models the platform hosts, verified against fal's current catalog. The image, editing, video, sound, and 3D/face shelves now name the long-tail specialists that make the garden worth touring: Bria's licensed background removal and SAM2 for matting, fal's own AuraSR and the CCSR/DRCT/ESRGAN upscaler rack, ByteDance SeedVR2 and Topaz for video restoration, the open TTS world (Kokoro, Chatterbox, Dia, Qwen3-TTS, MiniMax) plus Demucs stem separation, and Hunyuan3D, MuseTalk, and SadTalker on the 3D and face shelves.
VERSION 31July 2026
Added a new section to The Craft of Generation on the workflows opening up right now: the patterns working creators are reaching for as the tools mature. Covers first-and-last-frame keyframing, locking identity across shots with character sheets and identity LoRAs, instruction-based editing in chains, routing the best model to each step, stacking multiple references instead of describing in words, performing generation on a real-time canvas, and agentic graphs that run themselves.
VERSION 32July 2026
Currency notes on tooling and capital. ComfyUI's 2026 releases (int4/int8 quantization, native SeedVR2 upscaling, Depth Anything 3, SCAIL-2 character replacement, PixelDiT support) and ai-toolkit's new FLUX.2 training now appear in the field manual, and the economics chapter carries the mid-2026 funding surge in the inference layer (Together AI at 8.3B, Baseten near 13B, fal's AWS deal at 2.5M developers, Kuaishou weighing a Kling spin-off near 20B).
VERSION 33July 2026
Deepened the Specialized Models part, the thinnest in the book. The lipsync section now covers the widened field (MuseTalk realtime, SadTalker, the avatar platforms, Flawless visual dubbing); 3D covers Hunyuan 3D's PBR and retopology, TRELLIS LoRA support, and why 3D plugs straight into real pipelines; upscaling adds SeedVR2, AuraSR, and tiled diffusion with the reconstruction-versus-hallucination judgment; and editing notes the shift to chainable instruction editing.
VERSION 34July 2026
Merged the two back-to-back 'future' parts. 'Where the Field Is Headed' was a thin forecast chapter that largely duplicated the deeper 'The Research Frontier' right after it. The two are now one Research Frontier part: the deep technical sections first (distillation, consistency, long-form, world models, the state of the art), then the forward-looking forecasts and the visual roadmap, closing on the pace-of-innovation outlook. The most redundant stub (a 160-word 'better identity and consistency' that duplicated the 650-word deep section) was dropped, and a stale 'past week' framing on the HappyHorse passage was corrected.
VERSION 35July 2026
Structure and depth. Demoted the Inside fal deep-dive from a standalone part into a section of the Ecosystem field guide, so one vendor no longer sits at the same level as whole movements of the book; its shelves are preserved as subsections. Deepened the two thin analysis chapters: the comparative decision framework now walks the ordered questions a professional actually asks (still or moving, text or not, identity-locked or not, open-weight or not, quality last), the benchmarks section explains why average-preference scores miss the tail failures production work lives in, and The Stack's audio pillar spells out the speech, music, and foley layers and the coming pull back into the video box.
VERSION 36July 2026
Five new figures for dense concepts, each placed at the exact paragraph that explains it rather than at the top of the page: cross-attention (how the prompt reaches into every patch), ControlNet conditioning maps, the quantization precision ladder, first-and-last-frame keyframing, and identity locking with a reference stack and a fixed seed. Also repositioned the latent-diffusion figure from the top of the history chapter down to the section where latent diffusion is actually explained.
VERSION 37July 2026
Glossary pass. Added 23 more terms used across the book but undefined (tiled diffusion, virtual try-on, model routing, character sheet, PBR, retopology, matting, relighting, foley, voice cloning, node graph, previs, animatic, and more), bringing it to 132. Every glossary entry now links to the chapter that explains the term in depth, so the glossary works as a navigation hub, not just a lookup.
VERSION 38July 2026
Fixed a flow break the Foundations reorder introduced: 'From images to video' opened with 'everything so far applies to images, video is much harder,' but it sat just after the dataset-history chapter, which discusses video training data. Moved the video chapter up to directly follow 'Refinements on the basic recipe,' so it completes the image mechanics before the data chapters, and its opening reads true again. Also moved the 3D VAE figure down from the section top to the subsection that explains 3D VAEs.
VERSION 39July 2026
Two things. First, a full figure-placement pass: moved every figure that was stranded at the top of its section down to the paragraph that actually explains it, U-Net vs DiT, the denoising scrubber, MM-DiT, and the ComfyUI workflow graph (the latent-diffusion and 3D VAE figures were already relocated). Second, added eleven openly-licensed supplementary photos, self-hosted with attribution, on the pages where a real image helps: a neural network, an H100 accelerator, a silicon wafer, a data-center hall, a mixing console, a polygon mesh, a green-screen shoot, a color-grading panel, an editing bay, a robot arm, and a film set, each placed inline at the relevant paragraph.
VERSION 40July 2026
Expanded the ComfyUI Field Manual's agentic section into a full treatment of driving ComfyUI with MCP. New subsections cover what an agent can actually do through the Model Context Protocol, how to wire it up (Comfy Org's official hosted Comfy Cloud MCP beta, or a local server like artokun's control plane or joenorton's lightweight one, on top of ComfyUI's prompt endpoint and websocket), why it beats clicking nodes for batch and iteration, what it looks like in practice with real use cases and a case study, why MCP is a general ecosystem protocol rather than a ComfyUI feature, and the honest limits. Sourced to the Comfy Org blog and docs and the community repos.
VERSION 41July 2026
Deepened the Kling chapter into a full company-and-model profile: the Kuaishou business behind it and the mid-2026 spin-off (a roughly 2.8 billion dollar raise at about an 18 billion dollar valuation, Kuaishou keeping ~68 percent, Tencent, Alibaba, and Baidu backing, an IPO targeted later), an honest tech read (a confirmed diffusion-transformer stack via the Kling-Omni and Kling-Foley papers, with the popular 3D-VAE and full-attention specifics flagged as company claims), where Kling actually stands on the arenas and what it is best and worst at, how to get the most out of it (modes, access, pricing, prompting), and where it is headed. Corrected the unconfirmed 'Kling 3.0 Turbo' and the stale 'tied for the top spot' claim. Sourced to Kuaishou's filings and releases, the arXiv reports, the arenas, and API catalogs.
v42v42
Rebuilt the OpenAI Sora chapter into a full profile: the team and the DiT origin, the world-simulator thesis read honestly, what it was actually good and bad at, and the economics and legacy behind the 2026 shutdown.
v43v43
Deepened the Google Veo chapter with the organization and vertical-integration moat behind it, the A24 deal, and a practical guide to access, pricing, and prompting across Flow, the Gemini app, Vertex AI, and the API.
v44v44
Deepened the Runway chapter with the company's origin (NYU ITP, co-authoring Stable Diffusion, the deliberate NYC bet, the funding arc and the AI Film Festival), an honest strengths and weaknesses read, and a practical guide to plans, pricing, and when to reach for it.
v45v45
Added a Higgsfield deep dive to the ecosystem's aggregator layer: the cinematic camera-control differentiator and the model-wrapping bet, the steep 2026 revenue and funding ramp, and an honest read on the quality ceiling and the Forbes controversy.
v46v46
Folded the first weekly Field Report finding into the book: noted the July 23 launch of FLUX 3 in the Black Forest Labs chapter, marking the lab's jump from image generation to a unified multimodal and physical-AI model.
VERSION 46July 2026
Made the ComfyUI field manual readable by someone who has never opened it. Added plain-English analogies before the mechanics in the densest sections (the tool as a machine you run rather than a page you visit, the typed wires as plumbing that only fits when the gauges match, the six-node default graph as a darkroom assembly line, the denoise slider as a how-much-do-I-repaint dial, the API export as turning a hand-drawn recipe into a button, and the four ways to run as four ways to get a car), plus three new figures: the six typed data wires, the interface climbing from model to workflow to agent, and the four places ComfyUI can run. The graph and latent-pipeline diagrams now appear in the manual where the nodes are first walked through, so a newcomer can see why a graph, a latent, a sampler, and a VAE decode each exist.
VERSION 47July 2026
Expanded the figure engine. Ten new instrument-panel diagrams were added where a picture carries a dense idea better than prose: the three families of generators and why diffusion won, the dual text encoders that turn a prompt into a conditioning signal, full spatiotemporal attention in video, the ten dimensions of great generation as a radar, the anatomy of a senior prompt, how preference leaderboards resolve into ELO, the per-user unit economics stack, the five build-versus-buy decisions, why long video drifts, and the world-model loop. Existing plates were retuned so accented fills track the green signal color rather than the retired orange.
v48v48
Second video-model deep-dive batch: brought Seedance, Hailuo, Wan, Luma, Pika, and Grok Imagine to the same depth as Kling and Sora, with sourced business, architecture, honest strengths and weaknesses, and practical access guidance for each.
v49v49
Image-model deep-dive batch: brought FLUX, Google Nano Banana, Ideogram, Recraft, Midjourney, and Stability to Kling-and-Sora depth, with sourced business, honest architecture, practical access guidance, and the 2026 stories (FLUX 3 multimodal, the Nano Banana tier ladder, the Midjourney studio lawsuits, and Stability's collapse and reset).
v50v50
Second image-model batch, completing the by-company deep dives: OpenAI's native turn and the Ghibli moment, ByteDance Seedream's unified 4K engine, Alibaba Qwen-Image's open-then-closed drift, and Krea's honest aggregator-plus-lab nature, each with practical access guidance.
v51v51
Added a synthesis chapter reading the forces behind the models (distribution over quality, the open-then-closed drift, the training-data reckoning, the native-multimodal turn, and the capital divide), and did a light framing pass so a handful of passages lead with the dynamic and its consequence rather than a leaderboard position.
v52v52
Reverted the synthesis chapter that overreached into an unsupported distribution thesis. The focus stays on keeping the ecosystem coverage current, true, and accurate, not on grand competitive claims.
v53v53
Added a beginner orientation to the ComfyUI infrastructure chapter: why the node graph feels so complicated, why that complexity is the honest shape of the pipeline rather than a UI failure, and the reassurance that every workflow is the same six-node skeleton with small insertions.
v54v54
Threaded a concrete running example through the ComfyUI workflow walkthrough: one specific image (a cinematic astronaut portrait made with an SDXL checkpoint) followed node by node, with real prompts, settings, and model names, so the abstract graph is illustrated by an actual use case.
v55v55
Completed the training-data FAQ with the whose-computer-is-it-running-on spectrum: hosted UIs and APIs run on someone elses servers and can feed training, while running an open-weight model locally is private by construction because nothing leaves your machine.
v56v56
Audio and voice becomes a first-class part alongside image and video. This first drop adds the part, a foundations chapter on how audio generation actually works (codecs, TTS, music, sound, native audio), and a full ElevenLabs deep dive. More audio chapters follow.
v57v57
Audio part, drop two: the voice landscape beyond ElevenLabs (the real-time lane, the open lane, the workflow lane, and voice in filmmaking) plus full deep dives on the two music engines, Suno and Udio.
v58v58
Audio part, final drop: the music and sound landscape (Lyria, Stable Audio, MusicGen, Riffusion, Tencent, native audio in video) and a closing reckoning on licensing and voice likeness that ties the whole part together.
v59v59
Two figures for the audio foundations chapter: the neural codec as the audio VAE, and the two families of text-to-speech.
v60v60
Editorial pass, part one: audio is stitched into the cross-cutting chapters so the new pillar reads as equal to image and video, not a bolt-on. Added an audio landscape section to Comparative Analysis, an audio-quality section to Evaluating Quality, a directing-audio section to The Craft, audio unit economics, build-vs-buy, and a licensing moat to The Operator's Playbook, and a short audio-mechanics orientation to The Foundations.
v61v61
Editorial pass, part two: the AI Filmmaking part was consolidated from eleven overlapping short sections into six substantial ones and deepened with verified examples. The redundant how-the-pros-work sections merged into a single craft chapter, platforms merged with the agentic loop, films merged with where-AI-wins, and team-of-one merged with the authorship question. The sound step is deepened and tied explicitly to the audio pillar.
v62v62
Editorial pass, part three: depth. The Audio part gains a dedicated section on native audio inside the video models and the pressure it puts on standalone tools, a deepened Udio (fidelity, the UMG settlement, and the download-lockout backlash), and a fuller music and sound landscape (Lyria, the Riffusion-to-ProducerAI-to-Google arc, Stable Audio, AudioCraft licensing, and the Chinese open models). Specialized Models gains the current instruction-editing models (Nano Banana, Qwen-Image-Edit, FLUX Kontext, SeedEdit) and a rebuilt lipsync section from the Wav2Lip lineage through Sync, Hedra, HeyGen, Act-Two, and native video lip-sync.
v63v63
Editorial pass, part four: sequencing and naming. Renamed the audio part to Audio Models: Voice, Music, and Sound so the title no longer implies audio is only voice. Grouped the new audio-frontier section with the other frontier topics. Reframed the Research Frontier as the book's forward-looking capstone that now spans all three pillars, image, video, and audio, rather than reading as a detour back into methods.
v64v64
AI Filmmaking expanded roughly two to three times on deep, verified research. New sections on the craft of consistency and control, sound and post as a first-class stage tied to the audio pillar, where AI wins first by format, the studios and the money, festivals labor and the law, and the frontier with its honest limits. Existing sections deepened with the platform taxonomy, sourced generation-to-keep ratios, and verified case studies (PJ Ace, Trillo, Jacob Adler, Neural Viz, Metaphysic's Here, Critterz and the Sora shutdown as platform risk).
v65v65
Thirty-two new KEY INSIGHT callouts across the insight-starved but content-rich chapters (Foundations, Image, Video, LoRA, Specialized, Comparative, Evaluating Quality, the Research Frontier, the Ecosystem, the Stack, the ComfyUI manual, and Hollywood). Each is a non-obvious causal chain or reframe grounded in the chapter's own material, in the spirit of the alt-text insight, for example how export controls seeded open weights, why the diffusion objective labels itself, that video quietly solved image consistency, and that slop is a missing human step rather than a machine.
v66v66
Freshness and consistency pass after the recent expansion. Reordered the Research Frontier into a logical flow and merged out two redundant stubs (a thin world-models-as-simulation section and a thin real-time section, both fully covered by the world-models-and-interactive-generation chapter). Reconciled the Sora timeline so the economics and platform chapters no longer describe Sora as a currently-available product, matching the documented 2026 shutdown, and refreshed a stale April 2026 self-stamp to mid-2026.
v67v67
Five new instrument-panel figures for the dense new material: the seven-stage AI film pipeline, model-first vs orchestration-first platforms, the chain of ceilings (encoder, denoiser, VAE, plus the audio codec and 3D VAE), the two-consents rule for voice and likeness, and the inverted shooting ratio of generate-many-cull-hard.
v68v68
QA and currency pass. Fixed a batch of internal inconsistencies flagged by a consistency sweep (a stale FLUX 3 note, Wan wrongly called open-weight, the Sora shutdown date in the Hollywood chapter, ten-versus-fifteen craft dimensions, a mislabeled joint-audio-video first, fal model counts, Higgsfield founding year, and more), and added dated late-July field notes for fresh developments: the leaderboard reshuffle (GPT Image 2 and Gemini Omni Flash on top), the music-industry AI-labeling standard, the GEMA v. Suno verdict date, Udio's DRM enforcement, Netflix's InterPositive acquisition, and the Google DeepMind and A24 partnership.
v69v69
QA sweep, finished. Harmonized the last ambiguous date drifts (Nano Banana 2, Seedream 4.0, Krea 2) and a disputed Veo price by resolving the contradiction rather than guessing a value, updated two callouts that still hedged now-shipped models (Ideogram 4.0, Seedream 5.0 Pro), and removed three duplications introduced during the filmmaking expansion (the sound passage, the overgeneration ratios, and a repeated set of filmmaker case studies).
v70v70
Comprehension pass. Added a per-chapter TL;DR to every chapter (149) so you can take the gist before deciding how deep to go, extended active-recall coverage from 35 to 145 chapters (627 questions), expanded the preface to lay out how the book is built to help you read it, and added three figures (the four reliability failure modes, the Midjourney ratings flywheel, and the image-into-LLM shift).
v71v71
Glossary nearly doubled, from 132 to 257 terms. Added definitions for the jargon the recent expansions introduced: audio (neural codec, residual vector quantization, VALL-E, Voicebox, F5-TTS, native audio, stems, watermarking, SynthID, C2PA), filmmaking (performance transfer, digital double, de-aging, agentic production, the two-consents rule), business and policy (indemnification, licensed data, ELVIS Act, NO FAKES Act, model-first versus orchestration-first), and research and tooling (rectified flow, KSampler, Gaussian splatting, super-resolution, and more). Each new term also becomes a tap-to-define tooltip in the prose.
v72v72
Update sweep plus a deeper look at microdramas. Fresh notes: Midjourney V8.2 as the new default, Microsoft's first-party MAI-Image-2.5-Pro and MAI-Voice-2-Flash public preview, Meta's Muse Image and the consent debate it triggered, and the Delhi High Court's fair-dealing ruling as a third posture on the global copyright map. Deepened the microdrama coverage (its mobile-game economics, why the format fits current AI video, the localization-first path to AI, and distribution as the real moat) and expanded the ecosystem Microdramas directory with ReelShort, DramaBox, ShortMax, GoodShort, and more.
v73v73
Deeper coverage of the hybrid filmmaking process, the live-action production with AI slotted in at the edges that is where professional adoption actually lives. A new AI Filmmaking section walks the hybrid pipeline stage by stage (previs, on-set real-time de-aging on Here, the post and effects zone with Metaphysic, MARZ, Runway Aleph, Autodesk Flow Studio and Beeble, finishing with Topaz, and audio dubbing and visual dubbing), and makes the core argument explicit: the hybrid film is the only version of AI filmmaking that is simultaneously cheaper, copyrightable, and insurable, while a fully generated feature fails the human-authorship and clean-chain-of-title tests. It closes on the labor line, that SAG-AFTRA's consent regime blessed the performance while leaving effects and title artists exposed, and that training data is the unresolved fight.
v74v74
Rebuilt the hybrid filmmaking section as a much longer, instructional working guide. It now walks the process as how-to: deciding what to film versus generate (a new SHOOT THE MIDDLE, GENERATE THE ENDS insight), using previs to plan without committing, an on-set data-capture checklist (clean plates, lighting and lens reference, tracking, real-time reference, and consent capture), the post-production AI effects loop with tools mapped to tasks and an honest note that the finish is the job, finishing and localization, a consent, chain-of-title, and insurance clearance workflow, and a common-mistakes-plus-checklist close. Every concrete claim keeps an inline source link.
v75v75
Threaded inline source links through the book. Claims across 84 sections now carry a citation right where they are made (a model launch, a funding round, a benchmark, a paper, a court ruling, a dated event), each opening its source in a new tab, so you can go straight to the primary material. This complements the full provenance already listed on the updates page.
v76v76
Deepened the hybrid filmmaking guide again with the on-stage half and the finishing craft. Added the virtual-production pillar (the LED volume behind The Mandalorian and performance capture behind Avatar, with generative AI now feeding both, as in the Wonder Project and Luma venture Innovative Dreams), the mechanical reason the middle of a performance resists generation (no physical weight, texture too perfect, temporal degradation), and a new Techniques that sell the composite section covering the short-insert sandwich, the practical anchor plate, first-and-last-frame stitching, and a new insight that AI's real tell is that it is too clean, so you must degrade it back to reality. Grounded the cost pressure in FilmLA's 2025 production figures.
v77v77
Weekly refresh for the first days of August 2026. The law moved from pending to in force: a Munich court found Suno liable in GEMA's copyright suit, and the EU AI Act's Article 50 synthetic-content transparency rules took effect. Video kept absorbing audio, with MiniMax's omni-modal Hailuo H3 (2K video with native sound) and xAI's Grok Imagine 1.5 (references, voice, 1080p), while Google DeepMind's Lyria 3.5 pushed the music race and IFPI set chart-eligibility rules for AI music. Google also pulled its one-day-old Nano Banana feature in Google Earth, a reliability lesson.