Contents

Updates

A living document

The field moves.
This book moves with it.

A printed book about generative AI is out of date before the ink dries. This one is not printed. Each edition is reviewed against what actually shipped, sections that changed are marked in the contents, and everything below is the record of what moved and when.

Current edition: Version 78, August 11, 2026. Reviewed against the field roughly monthly. Every entry carries its sources; sections that changed are marked in the contents.

Version 78

August 11, 2026

  1. Upd

    Mid-August refresh: frontier video goes open, and AI music makes peace

    detail

    A weekly refresh for mid August 2026. On access: MiniMax open-sourced its 33B omni-modal H3 with native stereo audio and day-one ComfyUI support, Lightricks released open-weights LTX-2.5 that turns a still into a 10-second clip in about 6.8 seconds, and ByteDance opened the Seedance 2.5 API on Volcano Engine (native 30-second single-take video, up to 50 reference inputs). On standings: Gemini Omni Flash leads the Artificial Analysis Video Arena with MiniMax H3 at number two, while OpenAI's gpt-image-2 still tops the image arena and xAI's Grok Imagine Image 2.0 claims number two after an editing-first release. On rights: Suno signed a global licensing deal with BMG and committed to tamper-resistant watermarking and download caps, the clearest sign yet that AI music is trading its Wild West era for licensed catalogs and provenance. Model standings move weekly; treat any single ranking as a snapshot.

    Ch. 44Ch. 46Ch. 49Ch. 59Ch. 60Ch. 76

Version 77

August 3, 2026

  1. Upd

    Early August refresh: the law arrives, and video keeps absorbing audio

    detail

    A weekly refresh for early August 2026. On the law: the Munich Regional Court found Suno liable in GEMA's copyright suit (July 31), and the EU AI Act's Article 50 synthetic-content transparency rules took effect (August 2). On the models: MiniMax's omni-modal Hailuo H3 generates 2K video with native stereo sound in one pass, xAI's Grok Imagine 1.5 added references, voice, and 1080p, Google DeepMind's Lyria 3.5 pushed the music-generation race, and IFPI set chart-eligibility rules for AI music. Google also launched and then pulled a Nano Banana image feature in Google Earth within a day, a reliability lesson.

    Ch. 131Ch. 44Ch. 49Ch. 57Ch. 59Ch. 22

Version 76

July 31, 2026

  1. Upd

    Hybrid filmmaking: the stage, the composite techniques, and the finish

    structural

    The hybrid filmmaking guide gained its on-stage half and its finishing craft. New material covers virtual production (the LED volume that The Mandalorian made mainstream with ILM StageCraft) and performance capture (the Avatar lineage), both now fed by generative AI as in the Wonder Project and Luma venture Innovative Dreams; the mechanical reason a performance resists generation (no physical weight, texture too perfect, temporal degradation that worsens with shot length); and a Techniques that sell the composite section: keep the synthetic insert short and sandwich it between real anchors, grow a wide shot from a real anchor plate, bridge two real frames with a generated camera move, and degrade the clean AI image back toward reality with grain, halation, and an aggressive grade. Cost pressure is now grounded in FilmLA's 2025 figures (Los Angeles down to 19,694 shoot days, television down nearly sixty percent from its 2021 peak).

    Ch. 116

Version 75

July 31, 2026

  1. Upd

    Inline source links across the book

    detail

    Concrete claims throughout the book now link to their source inline, right where the claim is made, rather than only in the updates provenance list. Two hundred and forty five citations were threaded across eighty four sections, each anchored to a distinctive proper noun, dated event, or specific figure and opening the primary source in a new tab. The reader can now verify a model launch, a funding round, a benchmark result, a research paper, or a legal ruling at the point of reading.

Version 74

July 31, 2026

  1. Upd

    The hybrid filmmaking process, as a working guide

    structural

    The hybrid filmmaking section is now a full instructional guide to running a live-action production with AI at the edges. It covers the division of labor across the frame (shoot the performance, generate the front and back of the pipeline), previs as planning not commitment, an on-set capture checklist (clean plates, lighting and lens reference, tracking markers, a real-time reference as on Here, and consent capture on the day), the post-production effects loop with tools mapped to tasks (Metaphysic and MARZ for faces, Runway Aleph for in-plate edits, Autodesk Flow Studio for CG doubles, Beeble for relighting) and the reminder that the human finish is where the quality lives, finishing and localization (Topaz, ElevenLabs, Flawless), and a clearance workflow for consent, chain of title, and E and O insurance that makes a hybrid film ownable where a fully generated one is not. Closes on common mistakes and a prep-to-delivery checklist.

    Ch. 116

Version 73

July 31, 2026

  1. New

    The hybrid filmmaking pipeline, in depth

    structural

    A new AI Filmmaking section covers the hybrid process in depth: how generative AI enters a real live-action production at specific handoffs rather than making the whole film. It walks the stages (previs, on-set real-time de-aging on Zemeckis's Here, the deep post and effects zone with Metaphysic now in DNEG's Brahma, MARZ Vanity AI, Runway Aleph, Autodesk Flow Studio and Beeble, finishing with Topaz, and audio dubbing plus visual dubbing with ElevenLabs and Flawless), and argues that hybrid dominates because it is the only version that is cheaper, copyrightable, and insurable at once, where a fully generated film fails the human-authorship and errors-and-omissions chain-of-title tests. It ends on labor: SAG-AFTRA's consent regime made de-aging and voice refinement of an actor's own performance workable, the backlash concentrated on edge-of-frame uses that displaced effects and title artists, and training data remains the unresolved fight.

    Ch. 116

    Provenance

Version 72

July 25, 2026

  1. Upd

    Update sweep and a deeper microdrama chapter

    structural

    A freshness sweep added Midjourney V8.2, Microsoft's MAI-Image-2.5-Pro and MAI-Voice-2-Flash preview, Meta's Muse Image and its consent controversy, and the Delhi High Court's ANI v. OpenAI fair-dealing ruling. The microdrama vertical got a deeper treatment: it monetizes like a free-to-play mobile game, its short vertical format fits exactly what AI video does well, AI arrived first through dubbing and localization, and distribution scale (ReelShort and DramaBox reportedly hold about seventy percent of revenue) is the moat, not the model. The ecosystem Microdramas directory was expanded accordingly.

    Ch. 26Ch. 31Ch. 131Ch. 93

Version 68

July 25, 2026

  1. Upd

    Consistency fixes and a late-July currency pass

    structural

    A QA sweep reconciled internal contradictions introduced during the recent expansion (stale model hedges versus shipped notes, date and figure drift, Wan open-versus-closed, the Sora timeline in the Hollywood chapter). A freshness check against the leveled-up sources added dated notes: GPT Image 2 and Gemini Omni Flash now top the Artificial Analysis boards; the music industry launched a voluntary AI-Generated versus AI-Assisted labeling standard; Germany's GEMA v. Suno ruling is set for July 31, 2026; Udio added BuyDRM export enforcement; Netflix acquired Ben Affleck's InterPositive for about 587 million dollars; and Google DeepMind took a research stake in A24.

    Ch. 69Ch. 59Ch. 56Ch. 117Ch. 46

Version 64

July 25, 2026

  1. Upd

    AI Filmmaking, roughly tripled and built on deep research

    structural

    Expanded the flagship AI Filmmaking part with new sections on consistency and control (image-first, reference conditioning, identity LoRAs, camera and keyframe control, performance transfer), sound and post tied to the audio pillar, where AI wins by format (ads, music video, the microdrama boom, previs, VFX), the studios and the money (Promise, Asteria and Moonvalley, Primordial Soup, Staircase, Toonstar, Critterz), festivals labor and the law (SAG-AFTRA consent, the Copyright Office report, the Thaler Supreme Court denial), and the real-time world-model frontier. Deepened the platform taxonomy and the verified filmmaker case studies.

    Ch. 111Ch. 113Ch. 115Ch. 117Ch. 118Ch. 120

Version 63

July 25, 2026

  1. Upd

    Editorial polish: audio naming and the frontier reframed as capstone

    structural

    Closed the editorial pass with structure and naming. The audio part is renamed Audio Models: Voice, Music, and Sound so its title reflects that audio is voice, music, and sound design, not voice alone. The audio-frontier section now sits with the other frontier topics. And the Research Frontier is reframed as the book's forward-looking capstone spanning image, video, and audio, resolving the jump from the industry chapters back into research methods.

    Ch. 142

    Provenance

Version 62

July 25, 2026

  1. Upd

    Specialized Models deepened: instruction editing and lipsync

    structural

    Deepened two of the hottest specialized areas. Editing now covers the current instruction-based models (Gemini 2.5 Flash Image / Nano Banana, Qwen-Image-Edit, FLUX.1 Kontext, ByteDance SeedEdit and Seedream) and why surgical, multi-turn, consistent-character editing anchors real workflows. Lipsync was rebuilt from the Wav2Lip research lineage through Sync.so, Hedra Character-3, HeyGen, Runway Act-Two, and lip-sync now landing natively inside video models.

    Ch. 60Ch. 61

  2. Upd

    Audio part deepened: native audio, Udio, and the music landscape

    structural

    Moved the Audio part closer to parity with image and video. Added a standalone section on native audio in video (Veo 3, Sora 2, Seedance, Wan, Kling, Grok) and the up-market pressure it creates. Deepened Udio (its fidelity edge, the October 2025 UMG settlement, and the download-lockout backlash) and the music and sound landscape (Lyria across Google's ecosystem, Google's acquisition of Riffusion via ProducerAI, Stable Audio's licensed-data play, AudioCraft's non-commercial weights, and Tencent and ACE-Step on the open frontier).

    Ch. 58Ch. 56Ch. 57

Version 61

July 25, 2026

  1. Upd

    AI Filmmaking rebuilt: from eleven stubs to six deep chapters

    structural

    Consolidated and deepened the flagship AI Filmmaking part. New arc: vibe directing as the new posture, the platforms and the agentic production loop, how an AI film actually gets made (with sound as a first-class stage tied to the audio pillar), the craft of technique-volume-and-the-edit, the films and filmmakers and where AI wins first, and the team of one and the author's chair. Verified examples include the Kalshi Veo 3 ad, Neural Viz, OpenArt's Director and vibe directing, Runway Gen-4/Aleph/Act-Two, the Runway AI Film Festival and Total Pixel Space, Critterz, the Here de-aging work, and Aronofsky's Primordial Soup.

    Ch. 108Ch. 109Ch. 110Ch. 112Ch. 114Ch. 119

Version 60

July 25, 2026

  1. New

    Audio becomes first-class across the cross-cutting chapters

    structural

    Editorial pass stitching the audio pillar into the chapters that previously covered only image and video. Comparative Analysis now has an audio model landscape (voice latency, cloning, and the music structure and licensing axes). Evaluating Quality now has an audio quality section (naturalness, artifacts, consistency, controllability, intelligibility, provenance, weighted by use case). The Craft now covers directing voice, music, and sound, including the two-consents rule for cloning. The Operator's Playbook now models audio unit economics (speech meters on volume, music sells a cap), a hosted-vs-self-hosted decision, and the licensed-data-plus-indemnification moat. The Foundations gains a one-page audio-mechanics orientation.

    Ch. 66Ch. 87Ch. 107Ch. 09Ch. 129Ch. 142

Version 59

July 24, 2026

  1. Upd

    Figures for the audio foundations chapter

    detail

    Added two instrument-panel figures to the audio foundations chapter: the neural codec as the audio VAE (waveform to residual-vector-quantized tokens and back), and the two text-to-speech families (autoregressive codec language models versus flow matching).

    Ch. 52

Version 58

July 24, 2026

  1. New

    Audio part: the music landscape and the licensing reckoning

    structural

    Final audio drop, completing the part. Added the music and sound landscape (Lyria, Stable Audio, Meta AudioCraft, Riffusion, Tencent, sound effects and foley, and the native-audio-in-video bifurcation that pushes standalone tools up-market) and a closing reckoning chapter on licensing and voice likeness: the turn toward licensed data, the ELVIS and NO FAKES laws and SAG-AFTRA consent regime, provenance via watermarking and C2PA, and the through-line that the ability to generate a voice does not settle whether it is lawful or right.

    Ch. 57Ch. 59

Version 57

July 24, 2026

  1. New

    Audio part: the voice landscape, Suno, and Udio

    structural

    Second audio drop. Added the voice landscape beyond ElevenLabs (the real-time lane with Cartesia, OpenAI, Google, Hume, Sesame and the Scarlett Johansson Sky episode; the open lane with Kokoro, Chatterbox, Fish Speech and Microsoft declining to release VALL-E 2; the editing and enterprise lane; and voice in filmmaking, where speech-to-speech conversion beats text-to-speech), plus full deep dives on Suno and Udio, including their mirror-image settlement paths with the record labels.

    Ch. 54Ch. 55Ch. 56

Version 56

July 24, 2026

  1. New

    Audio and Voice becomes a first-class part (foundations + ElevenLabs)

    structural

    Added a new Audio and Voice part alongside image and video, opening with a foundations chapter on how audio generation works (neural codecs as the audio VAE, the two TTS families and voice cloning, why music is harder than speech, sound effects, and native audio in video) and a full ElevenLabs deep dive covering the company, the model line and product surface, getting the best output and holding a character voice across a film, and the likeness and consent problem. Voice, music, and the licensing reckoning follow in subsequent drops.

    Ch. 52Ch. 53

Version 55

July 24, 2026

  1. Upd

    Training-data FAQ: the local versus hosted distinction

    structural

    Answered the missing case in the will-my-prompts-train-the-next-model FAQ: hosted company UIs and APIs run on the providers servers and can log and train on your inputs, while running open-weight models locally through something like ComfyUI keeps everything on your machine, private by construction rather than by promise.

    Ch. 15

    Provenance

Version 54

July 24, 2026

  1. Upd

    ComfyUI walkthrough, now with a worked example

    structural

    Rewrote the ComfyUI workflow walkthrough to follow one concrete image (a cinematic astronaut in a wheat field, made with an SDXL checkpoint like Juggernaut XL) through every node, with specific prompts, a 25-step DPM++ 2M Karras KSampler setup, and named variations (image-to-image, LoRA, upscale, ControlNet, and swapping in a Wan or Hunyuan video loader), so the abstract graph is grounded in an actual use case.

    Ch. 14

    Provenance

Version 53

July 24, 2026

  1. Upd

    ComfyUI, for beginners: why it feels so complicated

    structural

    Added a subsection to the ComfyUI infrastructure chapter that meets a beginner's overwhelm head on, with analogies (the airplane cockpit, the microwave versus the professional kitchen, the graph as a language) and the core why: ComfyUI does not add complexity, it reveals the complexity that was always there, and every workflow is the same six-node skeleton with small insertions.

    Ch. 14

    Provenance

Version 50

July 24, 2026

  1. Upd

    Krea, in depth: aggregator, canvas, and model lab

    detail

    Deepened the Krea chapter with its honest dual nature as a real-time canvas and aggregation layer that also trains its own models, the flag that Krea 1 rode the FLUX ecosystem and Krea Realtime 14B distilled Wan while only Krea 2 is from scratch, its real-time and aesthetic-post-training edge, and practical access and workflow guidance.

    Ch. 30

  2. Upd

    Qwen-Image, in depth: the open bet and the closed drift

    detail

    Deepened the Alibaba Qwen-Image chapter with the 20B MMDiT on a frozen Qwen2.5-VL encoder, the Apache 2.0 open weights and standout Chinese text rendering, the honest 2026 drift as Qwen-Image-3.0 shipped closed with no weights or report, and practical open-weight, quantized-local, and API access guidance.

    Ch. 28

  3. Upd

    Seedream, in depth: the unified 4K engine

    detail

    Deepened the ByteDance Seedream chapter with the unified Seedream 4.0 (one diffusion transformer for generation, editing, and composition at 4K in about 1.8 seconds), the launch arena double-top over Nano Banana that later eroded, the secondary-sourced 4.5 and 5.0 flags, and practical access across Dreamina, Doubao, and BytePlus with the US caveat.

    Ch. 29

  4. Upd

    GPT-Image, in depth: the native turn and the Ghibli moment

    detail

    Deepened the OpenAI image chapter with the March 2025 shift from standalone DALL-E to native generation inside GPT-4o, the viral Studio Ghibli moment and the melting-GPUs demand, the gpt-image-1 through gpt-image-2 lineage now topping the arena, and practical access, token pricing, and the honest weaknesses (speed, the warm tint, moderation).

    Ch. 23

Version 49

July 24, 2026

  1. Upd

    Stability, in depth: the collapse, the license, and the reset

    detail

    Expanded the thin Stability chapter into a full profile: the SD3 license backlash that drove the community to FLUX, the two architectural eras (latent diffusion to MMDiT rectified flow), honest strengths and weaknesses, the near-bankruptcy and the Akkaraju, Parker, and Cameron reset, practical open-weight and licensing guidance, and the honest flag that the Stable Diffusion 4 reports are SEO blog spam.

    Ch. 27

  2. Upd

    Midjourney, in depth: the lawsuits and how to use it

    detail

    Deepened the Midjourney chapter with the Disney, Universal, and Warner Bros. copyright lawsuits and the training-data opacity behind them, practical access and parameter prompting guidance (--ar, --stylize, --sref, RAW), and the bootstrapped, profitable independence funding a long arc toward video, 3D, and world simulation.

    Ch. 26

  3. Upd

    Recraft, in depth: the design vertical

    detail

    Deepened the Recraft chapter with its design-and-brand positioning, the red panda stealth arena win with V3, the editable-vector and typography differentiators, the V4 and V4.1 design-taste rebuilds, and practical guidance on when to choose it over general image models.

    Ch. 25

  4. Upd

    Ideogram, in depth: the eroding typography moat

    detail

    Deepened the Ideogram chapter with the honest read that its text-rendering moat narrowed as DALL-E, Imagen, FLUX, Recraft, and GPT-native generation caught up, the gap between its own and neutral benchmark rankings, the unverified Ideogram 4.0 open-weight reports, and practical prompting for legible in-image text.

    Ch. 24

  5. Upd

    Nano Banana, in depth: the Gemini-native consolidation

    detail

    Deepened the Google image chapter with the consolidation onto the Gemini-native path (Imagen 4 retirement, migration to Nano Banana 2), the confusing Pro-versus-2 tier naming, the mid-2026 arena standing behind GPT Image 2, and practical guidance on the Pro, 2, and 2 Lite ladder plus the identity-consistency and conversational-editing strengths.

    Ch. 22

  6. Upd

    FLUX, in depth: the multimodal turn and how to use it

    detail

    Deepened the Black Forest Labs FLUX chapter with the 2026 multimodal turn (FLUX.2, the July FLUX 3 unified image, video, audio, and action model, the FLUX mimic robotics model and Audi test, and the 300-million-dollar Series B) plus practical open-versus-pro access and licensing guidance.

    Ch. 21

Version 48

July 24, 2026

  1. Upd

    Grok Imagine, in depth: access and the moderation problem

    detail

    Added practical access, pricing, and prompting guidance to the xAI Grok Imagine chapter, plus an honest, sourced account of the permissive-moderation stance and the nonconsensual-deepfake controversy and regulatory response that followed.

    Ch. 49

  2. Upd

    Pika, in depth: the consumer bet

    detail

    Deepened the Pika chapter with the Guo and Meng founding and their diffusion-research pedigree, the deliberate choice to own the fun, social, effects-driven corner of AI video rather than the frontier, honest strengths and weaknesses, and practical guidance on using the effects and ingredient features.

    Ch. 48

  3. Upd

    Luma, in depth: the world-model turn and the HUMAIN raise

    detail

    Deepened the Luma chapter with its pivot from video generation to unified world models (Ray3 reasoning, Uni-1 and Luma Agents), the 900-million-dollar HUMAIN Series C and the Project Halo supercluster, honest strengths and weaknesses, and practical access and HDR-pipeline guidance.

    Ch. 47

  4. Upd

    Wan, in depth: the open-core split

    detail

    Expanded the Alibaba Wan chapter with the crucial 2026 finding that its open weights stop at Wan 2.2 while 2.5, 2.6, and likely 2.7 are closed API-only products (an open-core split, not full openness), plus honest strengths and weaknesses and open-versus-hosted access guidance.

    Ch. 45

  5. Upd

    Hailuo, in depth: the honest read and the studio lawsuit

    detail

    Added an honest strengths-and-weaknesses read to the MiniMax Hailuo chapter (the stale mid-2025 arena ranking, the 456B parameter figure that is actually the M1 language model, and the Disney, Universal, and Warner lawsuit) plus practical access and director-camera prompting guidance.

    Ch. 44

  6. Upd

    Seedance, in depth: the four-name model and how to use it

    detail

    Added practical access guidance (Dreamina, Jimeng, Doubao, CapCut, BytePlus, and the US-availability caveat), the honest resolution-spec inconsistency, and the unified-multimodal direction including Seedance 2.0's native audio, to the ByteDance Seedance chapter.

    Ch. 46

Version 46

July 24, 2026

  1. Upd

    The ComfyUI field manual, made understandable

    detail

    Rewrote the densest ComfyUI passages for a general reader. Each hard section now opens with an everyday analogy before the machinery: the installed tool as a small power plant you run rather than a storefront you visit, the typed wires as plumbing that only connects when the gauges match, the six-node default workflow as a darkroom assembly line where the painter works on a compressed draft (which is why latents, samplers, and a VAE decode each exist), the denoise dial as how much of the canvas you repaint, the API export as turning a hand-drawn recipe into a callable button, and owning versus renting a GPU as owning versus renting a car. New figures render the six typed data wires, the model-to-workflow-to-agent climb, and the four ways to run, and the canonical graph and latent-pipeline diagrams now sit inside the manual itself.

    Ch. 70Ch. 71Ch. 72Ch. 73Ch. 77Ch. 79

    Provenance

  2. Upd

    FLUX 3: Black Forest Labs goes multimodal

    structural

    Noted the July 23 launch of FLUX 3, a unified image, video, audio, and action model with native synchronized audio and a companion robotics model (FLUX mimic), in the Black Forest Labs chapter. This is the first finding promoted out of the weekly Field Reports into the standing book.

    Ch. 21

Version 45

July 23, 2026

  1. New

    Higgsfield, in depth: the aggregator bet and its honest read

    detail

    Added a Higgsfield profile to the ecosystem aggregator layer: the cinematic camera-control differentiator, the DoP and Soul models wrapping frontier engines like Kling, Veo, and Sora, the steep 2026 revenue and funding ramp toward a reported 1.3 billion valuation, and an honest read on the wrapped-model quality ceiling and the February 2026 Forbes controversy over fake demos, deepfake content, and billing.

    Ch. 96

Version 44

July 23, 2026

  1. Upd

    Runway, in depth: the origin, the moat, and how to use it

    detail

    Deepened the Runway chapter with the company's NYU ITP founding and its co-authorship of Stable Diffusion, the funding arc and the AI Film Festival, an honest read on where its editing-first control beats raw generation and where it trails the frontier, and a practical guide to plans, pricing, and treating Runway as an editing and orchestration hub rather than a clip slot machine.

    Ch. 43

    Provenance

Version 43

July 23, 2026

  1. Upd

    Veo, in depth: the org, the moat, and getting the best out of it

    detail

    Deepened the Google Veo chapter with the DeepMind organization and the vertical-integration moat (in-house TPUs, the YouTube corpus, and distribution across Flow, Gemini, Vertex AI, and the API), the reported seventy-five million dollar A24 partnership, and a practical guide to the four access paths, per-second and credit pricing, and how to prompt for camera, lens, and native audio.

    Ch. 41

Version 42

July 23, 2026

  1. Upd

    Sora, in depth: the team, the world-simulator thesis, and the shutdown

    detail

    Rebuilt the OpenAI Sora chapter into a full profile: Tim Brooks and Bill Peebles and the DiT paper Sora was built on, the spacetime-patch world-simulator thesis read honestly against what actually emerged, Sora 2's native audio and Cameos, and the economics and cultural legacy behind the 2026 app shutdown and API sunset.

    Ch. 40

Version 41

July 23, 2026

  1. Upd

    Kling, in depth: the business, the tech, and where it stands

    detail

    Expanded the Kling chapter with Kuaishou's business and the July 2026 spin-off (a ~$2.8B raise near an $18B valuation, Kuaishou keeping ~68%, Tencent/Alibaba/Baidu backing), an honest architecture read from the Kling-Omni and Kling-Foley papers, its mid-2026 arena standing (~#6, behind Gemini Omni Flash, Seedance, and Wan), and practical guidance on modes, access, pricing, and prompting.

    Ch. 42

Version 40

July 23, 2026

  1. Upd

    Driving ComfyUI agentically with MCP

    detail

    Comfy Org shipped an official hosted Comfy Cloud MCP server (public beta, June 2026) that lets an AI assistant search, assemble, run, and iterate on ComfyUI workflows, alongside a growing set of local community MCP servers (Artokun's control plane, Joe Norton's lightweight server) that can also edit the live graph and install nodes and models.

    Ch. 77

Version 32

July 23, 2026

  1. Upd

    ComfyUI's 2026 releases: quantization, SeedVR2, and new conditioning

    detail

    ComfyUI added int4/int8 quantization, native SeedVR2 upscaling, Depth Anything 3, SCAIL-2 multi-reference character replacement, and PixelDiT support through mid-2026, while ai-toolkit added LoRA training for FLUX.2 and Z-Image.

    Ch. 78

    Provenance

Version 26

July 23, 2026

  1. Upd

    AI copyright suits advance; EU AI Act enforcement begins August 2

    detail

    A judge let the studios' suit against MiniMax's Hailuo proceed (May 26), the New York Times sought sanctions against OpenAI over discovery (July), and the EU AI Act's enforcement powers over general-purpose model providers take effect August 2, 2026 with fines up to 15M euros or 3% of global turnover.

    Ch. 10Ch. 135

  2. Upd

    Studios move from lawsuits to equity: A24, Lionsgate, Bertelsmann, Getty

    detail

    Mid-2026 saw rights-holders take stakes and sign licensing deals: Google DeepMind invested $75M in A24 to co-develop Veo tools, Lionsgate took an equity stake in Runway (with a Bertelsmann partnership following), and Getty Images signed a display-only deal with OpenAI for ChatGPT.

    Ch. 123

Version 24

July 23, 2026

  1. Upd

    Google's Nano Banana 2 rolls out in three tiers

    detail

    Google shipped Nano Banana 2 (Gemini 3.1 Flash Image) in February 2026 across Gemini, Search, and Ads, in Lite, standard, and Pro tiers, with several variants ranking in the image-arena top ten.

    Ch. 22

    Provenance

  2. Upd

    Ideogram 4.0 ships as the lab's first open-weight model

    detail

    Ideogram released Ideogram 4.0 on June 3, 2026: a 9.3B single-stream diffusion Transformer with a Qwen3-VL text encoder, native 2K resolution, and JSON prompting, its first model with open weights alongside API access.

    Ch. 24

    Provenance

  3. Upd

    Black Forest Labs launches FLUX 3, its first multimodal model

    detail

    FLUX 3 launched July 23, 2026 as a multimodal frontier model generating images plus video with synchronized audio (clips up to roughly 20 seconds), in limited early access, extending Black Forest Labs beyond still images.

    Ch. 21

    Provenance

  4. Upd

    GPT Image 2 tops the image arena

    detail

    OpenAI released GPT Image 2 (ChatGPT Images 2.0) on April 21, 2026, retiring DALL-E 3 and GPT-Image-1.5 with near-perfect text rendering, higher resolution, and a thinking mode. It currently ranks first on the Artificial Analysis text-to-image arena.

    Ch. 23Ch. 69

    Provenance

Version 23

July 23, 2026

  1. Upd

    Real-time world models draw funding and open releases

    detail

    World models became the fastest-moving frontier category in 2026: Reactor emerged from stealth with a $59M round to build a developer platform for real-time AI worlds, alongside World Labs' Marble, Google DeepMind's Genie 3, and Tencent's open real-time HY-World.

    Ch. 97

  2. Upd

    The inference layer's 2026 funding surge

    detail

    The compute layer that serves generative media drew huge rounds in mid-2026: fal named AWS its preferred cloud at roughly 2.5 million developers, Baseten raised a $1.5B Series F near a $13B valuation, and Together AI raised $800M at $8.3B, all on rapidly growing inference demand.

    Ch. 99

Version 22

July 23, 2026

  1. Upd

    Kuaishou weighs a Kling spin-off near a $20B valuation

    detail

    Kuaishou disclosed in a May 12 exchange filing that it is evaluating a restructuring of the Kling AI business, with reports pointing to a pre-IPO round near a twenty-billion-dollar valuation as Kling's revenue run rate climbed toward five hundred million dollars.

    Ch. 42

  2. Upd

    Luma's Ray3.2 adds frame-level control and EXR output

    detail

    Luma released Ray3.2 (June 9), adding frame-by-frame directability and HDR generation with paired EXR outputs for professional pipelines, followed by 'Skills' repeatable agent workflows (June 16).

    Ch. 47

    Provenance

  3. Upd

    Seedance 2.5 claims a thirty-second single take

    detail

    ByteDance unveiled Seedance 2.5 at its Volcano Engine FORCE conference (June 23), claiming native single-shot video up to thirty seconds and up to roughly fifty reference inputs, with public API access via BytePlus opening July 16.

    Ch. 46

  4. Upd

    Runway's 2026 turn: Aleph 2.0, a Media Router, and studio deals

    detail

    Runway shipped Aleph 2.0 and Edit Studio (May 21), then a Media Router that selects across image, video, and audio models (July 23), while taking a Lionsgate equity stake (June 11) and announcing a Bertelsmann partnership (July 1). The pattern is a bet on owning workflow and distribution rather than only the base model.

    Ch. 43

  5. Upd

    Google moves video into Gemini Omni; Omni Flash tops the arenas

    detail

    Google shipped no Veo 4 at I/O 2026 and instead launched Gemini Omni Flash (June 30), a fast video generation and conversational-editing model that now ranks first on both the text-to-video and image-to-video arenas. Veo 3.1 remains the standalone flagship.

    Ch. 41Ch. 51

Version 21

July 23, 2026

  1. New

    ComfyUI, from setup to senior

    structural

    The ComfyUI manual gains a mastery arc: where it runs (local vs rented GPUs, RunPod, serverless, Comfy Cloud), the pitfalls that cost beginners a weekend, the 80/20 path to getting good, how professionals actually build with it, and what 'senior at ComfyUI' means.

    Ch. 79Ch. 80Ch. 81Ch. 82Ch. 83

    Provenance

  2. New

    In the studio: how the best AI creators actually work

    structural

    Two chapters on process: deep profiles of how PJ Accetturo, Neural Viz, Dave Clark, Paul Trillo, and the Dor Brothers actually make their work, and the patterns that recur across all of them (the common stack, generate-then-cull ratios, model-mixing, and editing discipline).

    Ch. 112

    Provenance

Version 20

July 22, 2026

  1. New

    Inside fal: a tour of the model garden

    structural

    A new chapter uses fal, the developer inference platform hosting 1,000+ generative models, as a vantage point on the whole medium: a guided tour of the interesting and less-mainstream models by category, profiling Bria (licensed-data image), Reve, Sonilo (video-to-music), Decart (real-time world models), and Lightricks' open LTX-2, among many others.

    Ch. 101

    Provenance

    • fal Series C announcementT1verified 2026.07.22The Generative Media Platform for Developers; in the last 12 months our revenue grew 60x.
    • fal (models and platform)T1verified 2026.07.22The world's best generative image, video, and audio models, all in one place; 1,000+ production-ready models.
    • Bria (licensed-data visual AI)T1verified 2026.07.22Enterprise visual AI trained on 100% licensed data with full IP indemnification and an attribution engine.
    • Decart publications (real-time world models)T1verified 2026.07.22The infrastructure and models that make AI run at the speed of reality; Oasis and MirageLSD.
    • Lightricks LTX-2T1verified 2026.07.22Open-weights synchronized audio and video at native 4K, runnable on consumer GPUs.

Version 19

July 22, 2026

  1. Upd

    Freshness pass: current model versions and arena standings

    detail

    Verified the mid-2026 landscape against the live leaderboards and updated the stragglers: Midjourney V8.1 is now the default, ByteDance shipped Seedream 5.0 Pro, and the current arena order (GPT Image 2 on top, Reve 2.1 and Microsoft MAI-Image-2.5 behind; Gemini Omni Flash leading video) is confirmed.

    Ch. 26Ch. 29Ch. 32Ch. 51

    Provenance

Version 18

July 22, 2026

  1. New

    A ComfyUI field manual

    structural

    A thorough hands-on manual for ComfyUI in The Machine: install and setup, reading the node graph, the default text-to-image workflow node by node, the core recipes (img2img, inpaint, outpaint, upscale, LoRA, ControlNet, IPAdapter), native video and audio, the Manager and registry, the headless/agentic API, and practical memory and reproducibility.

    Ch. 70Ch. 72Ch. 73Ch. 74Ch. 75Ch. 77Ch. 78

    Provenance

    • ComfyUI (GitHub)T1verified 2026.07.22The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
    • ComfyUI documentationT1verified 2026.07.22Official node reference and tutorials, including the text-to-image graph and KSampler parameters.
    • Comfy Org raises $17MT1verified 2026.07.22Open source must win. If a proprietary service dominates, creativity loses.
  2. New

    AI filmmaking, deeper: platform mechanics, where it wins, and the team of one

    structural

    Three more chapters on the craft: how the agentic loop actually runs inside Director, Flow, LTX Studio, and Showrunner; where AI filmmaking is already winning (ads, music videos, microdramas, previs, VFX) and where it still loses; and the real team shape and economics of an AI production, from Critterz to the solo creator.

    Ch. 109Ch. 114Ch. 119

    Provenance

Version 17

July 22, 2026

  1. New

    The book, restructured into two movements, and a new Ecosystem chapter

    structural

    The book is now organized as two movements, The Machine (how generative models work) and The Medium (what is being built on them). A new chapter walks the entire creative stack in prose, phase by phase, naming the companies doing the work in each box.

    Ch. 95Ch. 96Ch. 97Ch. 98Ch. 99Ch. 100

    Provenance

    • Machine Cinema (market map)T2verified 2026.07.22A market map of the AI creative ecosystem across development, production, post, inference, and distribution.
  2. New

    AI filmmaking, in depth: the pipeline, the techniques, and the films

    structural

    The AI Filmmaking part gains three craft chapters: how an AI film actually gets made end to end, the techniques that separate finished work from slop, and the real films and filmmakers driving the moment (PJ Accetturo, Neural Viz, Curious Refuge, Runway's festival, and the Sora-made feature Critterz).

    Ch. 110Ch. 112Ch. 114

    Provenance

Version 16

July 22, 2026

  1. New

    The creator scene, the memes, and the slop fight

    structural

    The native AI-creator canon (Will Smith spaghetti, the Ghibli trend, Neural Viz, PJ Ace, Curious Refuge, Runway's AIFF) and the two backlash fights, over slop (Willison's definition) and over consent (Glaze/Nightshade, Tilly Norwood, ScarJo vs Sky, likeness vaults, the jobs debate).

    Ch. 124Ch. 125

    Provenance

  2. New

    The deals: Lionsgate, Netflix, Disney-OpenAI, the Sphere

    structural

    What the studios actually did: Lionsgate-Runway's underdelivery and restructuring, Netflix's ~300 AI-assisted 2026 titles, the Disney-OpenAI Sora deal and its March 2026 collapse plus Disney's lawsuits, and the Sphere's $370M+ AI Wizard of Oz.

    Ch. 123

    Provenance

  3. New

    The auteurs and the guilds weigh in

    structural

    A new Industry act opens with the directors' split (Scorsese joining Black Forest Labs, Cameron, Lucas, Jackson vs del Toro, Villeneuve, Spielberg, Nolan as DGA president) and the institutional rules (WGA/SAG-AFTRA/DGA deals, the Academy's AI eligibility rules, the White House letter).

    Ch. 121Ch. 122

    Provenance

Version 15

July 22, 2026

  1. New

    AI filmmaking: vibe directing and agentic production

    structural

    A new part covers the turn from generating clips to directing films: vibe directing (OpenArt's June 2026 coinage, the film analogue of vibe coding), the agentic platforms led by Runway and OpenArt, the broader movement, and the authorship debate that dominated Cannes 2026.

    Ch. 108Ch. 109Ch. 119

    Provenance

  2. Upd

    The book gets an act structure and a new title

    structural

    The contents are regrouped into named acts (Foundations, The Big Picture, The Engines, The Craft, The Business, The Frontier, Reference), the glossary moves to the back, and the title broadens to reflect that the story is now the whole stack and AI filmmaking, not only the models.

    Provenance

    • Machine Cinema market mapT2verified 2026.07.22The ecosystem view that motivated regrouping the book around the whole stack.

Version 14

July 22, 2026

  1. New

    Realtime and interactive generation

    structural

    A different mode from batch: realtime canvases (Krea, Decart) and world models generate as fast as you move, collapsing the pipeline and folding finishing and distribution into the moment of generation.

    Ch. 94

    Provenance

    • Decart (Mirage realtime video)T1verified 2026.07.22Sub-40ms/frame streaming diffusion for interactive, realtime video.
    • KreaT1verified 2026.07.22Realtime generation canvas that renders as fast as the user works.
  2. New

    Distribution and the microdrama economy

    structural

    Distribution is where the money lands. The vertical microdrama market (~$3.8B in-app in 2025, forecast to more than double in 2026) rivals Netflix for US mobile time; today it is mostly live action with AI as accelerant, but it is the format most exposed to full generation, with AI-native apps like Popshort and studios like Fairground emerging.

    Ch. 93

    Provenance

  3. New

    The audio pillar

    structural

    Audio (voice, dubbing, lip-sync, music, SFX) is nearly half the production phase and had been underweighted; a video pipeline without an audio plan is half a pipeline.

    Ch. 92

    Provenance

    • ElevenLabsT1verified 2026.07.22Leading AI voice, dubbing, music, and sound-effects generation.
    • Machine Cinema market mapT2verified 2026.07.22Voice, dubbing, lip-sync, music, and SFX are distinct first-class categories.
  4. New

    The production pipeline as a chain

    structural

    A finished AI-native piece is a chain of specialized jobs (write, board, generate, hold consistency, voice, lip-sync, edit, grade, mark, distribute); generation is one link and not the hardest.

    Ch. 91

    Provenance

    • Machine Cinema market mapT2verified 2026.07.22Each phase of the stack is a category because each is a distinct production job.
  5. New

    The workflow and aggregator layer

    structural

    Creators increasingly work in multi-model canvases (Krea, Higgsfield, Flora, ComfyUI) rather than going model-direct; the workflow, not the model, owns the user, and agents are starting to drive the workflow.

    Ch. 90

    Provenance

  6. New

    The Stack: the book gets a new spine

    structural

    A new anchoring part reframes the field from the point of view of the whole creative stack rather than individual models: as generation commoditizes, value migrates up to workflows and down to distribution. Concrete backing: the application layer now captures the majority of enterprise gen-AI spend, and app-layer players like Higgsfield reached a ~$500M run rate in about a year.

    Ch. 88Ch. 89

    Provenance

Version 13

July 22, 2026

  1. New

    When generation learned to reason

    structural

    A new Foundations chapter on reasoning-augmented image generation: GPT-Image-2's Thinking mode plans, web-searches, and self-checks before rendering (architecture undisclosed), alongside a genuine 2026 research revival of autoregressive image models.

    Ch. 18

    Provenance

  2. New

    When video learned to talk

    structural

    A new Foundations chapter on native audio-visual generation: Veo, Kling 3.0's Omni Native Audio, and open-weight LTX-2 generate synchronized sound in the same pass as the picture; open models like Foley-Omni handle video-to-audio.

    Ch. 19

    Provenance

  3. New

    Provenance, watermarking, and the August 2026 deadline

    structural

    A new Operator's Playbook section on the EU AI Act Article 50 transparency obligations that apply from August 2, 2026 (deepfake disclosure and machine-readable marking), and the C2PA plus SynthID standards converging to meet them.

    Ch. 135

    Provenance

  4. New

    World models leave the lab

    structural

    A new section promoting world models from research to product: DeepMind's Genie 3 (real-time interactive worlds, ~1-minute memory, research prototype) and World Labs' Marble (a commercial paid product that exports Gaussian splats and meshes).

    Ch. 141

    Provenance

  5. Upd

    Agents drive the ComfyUI graph

    structural

    Chapter 7 gains a subsection on Comfy MCP (launched June 30, 2026), the official bridge that lets AI agents assemble and run ComfyUI workflows on Comfy Cloud GPUs, plus in-canvas copilots.

    Ch. 14

    Provenance

    • Comfy MCPT1verified 2026.07.22Comfy Org launched Comfy MCP on June 30, 2026; agents run workflows on Comfy Cloud GPUs.
  6. Upd

    Blackwell economics and current pricing

    detail

    The economics chapter gains a hardware-and-pricing subsection: B200/GB200 rental ranges, Blackwell's inference-cost cut over Hopper, and current API reference prices (FLUX.2 from official docs; per-second video prices marked approximate).

    Ch. 13

    Provenance

Version 12

July 22, 2026

  1. Upd

    ComfyUI deep dive: samplers, steps, and CFG scale

    structural

    Chapter 7 gains a full treatment of the KSampler: what a denoising step does and where the returns flatten, how CFG scale trades prompt adherence against image quality, and how to pick a sampler and scheduler. The through-line is that settings must match the model class.

    Ch. 14

    Provenance

  2. Upd

    OpenAI exits consumer video; Sora API sunsets September 24

    detail

    The Sora app went dark in April and the API will be discontinued September 24, 2026, amid roughly a million dollars a day in cost, a dead Disney deal, and copyright problems. Sora continues only as internal world-model research.

    Ch. 40Ch. 51

    Provenance

    • OpenAI HelpT1verified 2026.07.22OpenAI: Sora app discontinued Apr 26, 2026; Sora API to be discontinued Sep 24, 2026.
    • The DecoderT3verified 2026.07.22
  3. Upd

    Google rebrands video as Gemini Omni

    detail

    Rather than a Veo 4, Google I/O introduced Gemini Omni, an any-to-any multimodal family. Gemini Omni Flash shipped the same day and sat at #1 across the major video arenas by July.

    Ch. 41Ch. 51

    Provenance

  4. New

    The Chinese video surge: Kling 3.0, Seedance 2.5, HappyHorse

    structural

    Kling 3.0 Turbo added multi-shot prompting and native audio as Kuaishou raised at a reported ~$18B valuation; ByteDance's Seedance 2.5 generates native single-pass 30-second 4K clips; Alibaba's HappyHorse topped the open video arenas with single-pass joint audio and video.

    Ch. 42Ch. 46Ch. 51

    Provenance

  5. New

    Meta enters image generation with Muse

    structural

    Meta Superintelligence Labs launched Muse Image, its first image model, an agentic system that plans a layout before rendering, shipping free inside the Meta AI app, Instagram, and WhatsApp.

    Ch. 32

    Provenance

  6. New

    The open-weights wave: Ideogram 4.0, Krea 2, HiDream-O1

    structural

    Three serious open-weight image models shipped in five weeks, closing the gap on the closed flagships: Ideogram 4.0 with structured-JSON prompting, Krea 2 with a two-second Turbo variant, and the MIT-licensed pixel-native HiDream-O1.

    Ch. 32

    Provenance

  7. Upd

    Google retires Imagen for the Gemini image line

    detail

    Google began sunsetting the Imagen family, with API endpoints shutting down through August 2026 and the migration path pointing to Nano Banana. Nano Banana 2 Lite became the fastest, cheapest tier at ~4s and ~3 cents per image.

    Ch. 22Ch. 32

    Provenance

  8. Upd

    The tier below Google and OpenAI is a race

    detail

    Recraft V4.1 briefly led outside Google and OpenAI in May, but by late July had been passed by Reve 2.1 and Microsoft's MAI-Image-2.5 on the Artificial Analysis arena.

    Ch. 32

    Provenance

  9. New

    New frontier entrants: Reve 2.1, MAI-Image-2.5, Grok Imagine 1.5

    detail

    Reve 2.1 and Microsoft's MAI-Image-2.5 now rank #2 and #3 on the Artificial Analysis image arena, and xAI's Grok Imagine Video 1.5 (May 31) briefly topped the image-to-video arena with native audio. All three were absent from earlier editions.

    Ch. 32Ch. 51

    Provenance