Updates
A living document
The field moves.
This book moves with it.
A printed book about generative AI is out of date before the ink dries. This one is not printed. Each edition is reviewed against what actually shipped, sections that changed are marked in the contents, and everything below is the record of what moved and when.
Current edition: Version 78, August 11, 2026. Reviewed against the field roughly monthly. Every entry carries its sources; sections that changed are marked in the contents.
Version 78
August 11, 2026
Upd
Mid-August refresh: frontier video goes open, and AI music makes peace
detailA weekly refresh for mid August 2026. On access: MiniMax open-sourced its 33B omni-modal H3 with native stereo audio and day-one ComfyUI support, Lightricks released open-weights LTX-2.5 that turns a still into a 10-second clip in about 6.8 seconds, and ByteDance opened the Seedance 2.5 API on Volcano Engine (native 30-second single-take video, up to 50 reference inputs). On standings: Gemini Omni Flash leads the Artificial Analysis Video Arena with MiniMax H3 at number two, while OpenAI's gpt-image-2 still tops the image arena and xAI's Grok Imagine Image 2.0 claims number two after an editing-first release. On rights: Suno signed a global licensing deal with BMG and committed to tamper-resistant watermarking and download caps, the clearest sign yet that AI music is trading its Wild West era for licensed catalogs and provenance. Model standings move weekly; treat any single ranking as a snapshot.
Ch. 44→Ch. 46→Ch. 49→Ch. 59→Ch. 60→Ch. 76→
Provenance
- Open General Intelligence: MiniMax H3 Is Now Open SourceT1verified 2026.08.11
- Lightricks/LTX-2.5 (Hugging Face)T1verified 2026.08.11
- ByteDance Seedance 2.5 API Goes Live - 30-Second Single-Shot Clips, 50 Reference Inputs, and 3D Camera BlockoutsT2verified 2026.08.11
- Suno and BMG Reach Licensing Deal for AI Music Model (Billboard)T2verified 2026.08.11
- Artificial Analysis — Text-to-Video LeaderboardT2verified 2026.08.11
- xAI Ships Grok Imagine Image 2.0 With Precise Editing and a Top Arena Ranking (Unite.AI, 2026-08-08)T2verified 2026.08.11
Version 77
August 3, 2026
Upd
Early August refresh: the law arrives, and video keeps absorbing audio
detailA weekly refresh for early August 2026. On the law: the Munich Regional Court found Suno liable in GEMA's copyright suit (July 31), and the EU AI Act's Article 50 synthetic-content transparency rules took effect (August 2). On the models: MiniMax's omni-modal Hailuo H3 generates 2K video with native stereo sound in one pass, xAI's Grok Imagine 1.5 added references, voice, and 1080p, Google DeepMind's Lyria 3.5 pushed the music-generation race, and IFPI set chart-eligibility rules for AI music. Google also launched and then pulled a Nano Banana image feature in Google Earth within a day, a reliability lesson.
Ch. 131→Ch. 44→Ch. 49→Ch. 57→Ch. 59→Ch. 22→
Provenance
- GEMA notches a second transatlantic AI copyright win in Germany (Reed Smith)T2verified 2026.08.03
- Guidelines on transparency obligations for AI-generated content (European Commission)T1verified 2026.08.03
- MiniMax H3 (official blog)T1verified 2026.08.03
- Imagine Video 1.5 with References (xAI news)T1verified 2026.08.03
- Lyria (Google DeepMind)T1verified 2026.08.03
- IFPI Global Principles for AI Recordings in Official Charts (IFPI)T1verified 2026.08.03
- Google nixes its Earth AI feature one day after launch (TechCrunch)T1verified 2026.08.03
Version 76
July 31, 2026
Upd
Hybrid filmmaking: the stage, the composite techniques, and the finish
structuralThe hybrid filmmaking guide gained its on-stage half and its finishing craft. New material covers virtual production (the LED volume that The Mandalorian made mainstream with ILM StageCraft) and performance capture (the Avatar lineage), both now fed by generative AI as in the Wonder Project and Luma venture Innovative Dreams; the mechanical reason a performance resists generation (no physical weight, texture too perfect, temporal degradation that worsens with shot length); and a Techniques that sell the composite section: keep the synthetic insert short and sandwich it between real anchors, grow a wide shot from a real anchor plate, bridge two real frames with a generated camera move, and degrade the clean AI image back toward reality with grain, halation, and an aggressive grade. Cost pressure is now grounded in FilmLA's 2025 figures (Los Angeles down to 19,694 shoot days, television down nearly sixty percent from its 2021 peak).
Provenance
- The Mandalorian (Season 1), StageCraft LED volume (Industrial Light and Magic)T1verified 2026.07.31
- How Avatar Redefined Motion Capture (No Film School)T2verified 2026.07.31
- Wonder Project and Luma Launch Innovative Dreams (Luma)T1verified 2026.07.31
- California Film and TV Tax Credit amid 2025 Production Losses (FilmLA)T1verified 2026.07.31
- FilmLA 2024 On-Location Production Report (FilmLA)T1verified 2026.07.31
- Dehancer film emulation, grain and halation (Dehancer)T1verified 2026.07.31
Version 75
July 31, 2026
Upd
Inline source links across the book
detailConcrete claims throughout the book now link to their source inline, right where the claim is made, rather than only in the updates provenance list. Two hundred and forty five citations were threaded across eighty four sections, each anchored to a distinctive proper noun, dated event, or specific figure and opening the primary source in a new tab. The reader can now verify a model launch, a funding round, a benchmark result, a research paper, or a legal ruling at the point of reading.
Provenance
- Video generation models as world simulators (OpenAI)T1verified 2026.07.31
- Scalable Diffusion Models with Transformers (arXiv)T1verified 2026.07.31
- Copyright and Artificial Intelligence, Part 2 (US Copyright Office)T1verified 2026.07.31
Version 74
July 31, 2026
Upd
The hybrid filmmaking process, as a working guide
structuralThe hybrid filmmaking section is now a full instructional guide to running a live-action production with AI at the edges. It covers the division of labor across the frame (shoot the performance, generate the front and back of the pipeline), previs as planning not commitment, an on-set capture checklist (clean plates, lighting and lens reference, tracking markers, a real-time reference as on Here, and consent capture on the day), the post-production effects loop with tools mapped to tasks (Metaphysic and MARZ for faces, Runway Aleph for in-plate edits, Autodesk Flow Studio for CG doubles, Beeble for relighting) and the reminder that the human finish is where the quality lives, finishing and localization (Topaz, ElevenLabs, Flawless), and a clearance workflow for consent, chain of title, and E and O insurance that makes a hybrid film ownable where a fully generated one is not. Closes on common mistakes and a prep-to-delivery checklist.
Provenance
- About 300 Netflix Titles Used Generative AI This Year (Variety, Q2 2026 earnings)T1verified 2026.07.31
- Tom Hanks, Robin Wright to Be De-aged Using Metaphysic AI Tool (The Hollywood Reporter)T1verified 2026.07.31
- Runway Partners with Lionsgate in First-of-its-Kind AI Collaboration (Lionsgate)T1verified 2026.07.31
- Runway Aleph: AI Edits Real Footage (CineD)T2verified 2026.07.31
- Digital Replicas 101 (SAG-AFTRA)T1verified 2026.07.31
- Copyright and Artificial Intelligence, Part 2: Copyrightability (US Copyright Office)T1verified 2026.07.31
- Introducing Dubbing v2 (ElevenLabs)T1verified 2026.07.31
- Will AI Influence the Oscar Race Amid The Brutalist Backlash? (Variety)T1verified 2026.07.31
Version 73
July 31, 2026
New
The hybrid filmmaking pipeline, in depth
structuralA new AI Filmmaking section covers the hybrid process in depth: how generative AI enters a real live-action production at specific handoffs rather than making the whole film. It walks the stages (previs, on-set real-time de-aging on Zemeckis's Here, the deep post and effects zone with Metaphysic now in DNEG's Brahma, MARZ Vanity AI, Runway Aleph, Autodesk Flow Studio and Beeble, finishing with Topaz, and audio dubbing plus visual dubbing with ElevenLabs and Flawless), and argues that hybrid dominates because it is the only version that is cheaper, copyrightable, and insurable at once, where a fully generated film fails the human-authorship and errors-and-omissions chain-of-title tests. It ends on labor: SAG-AFTRA's consent regime made de-aging and voice refinement of an actor's own performance workable, the backlash concentrated on edge-of-frame uses that displaced effects and title artists, and training data remains the unresolved fight.
Provenance
- Tom Hanks, Robin Wright to Be De-aged Using Metaphysic AI Tool (The Hollywood Reporter)T1verified 2026.07.31
- Metaphysic De-Ages Robert Zemeckis' Here Via Generative AI (Animation World Network)T2verified 2026.07.31
- Runway Aleph: AI Edits Real Footage with Camera Angles, Object Removal, and Relighting (CineD)T2verified 2026.07.31
- Runway Partners with Lionsgate in First-of-its-Kind AI Collaboration (Lionsgate)T1verified 2026.07.31
- Will AI Influence the Oscar Race Amid The Brutalist Backlash? (Variety)T1verified 2026.07.31
- Introducing Dubbing v2 (ElevenLabs)T1verified 2026.07.31
- About 300 Netflix Titles Used Generative AI This Year (Variety, Q2 2026 earnings)T1verified 2026.07.31
- Copyright and Artificial Intelligence, Part 2: Copyrightability (US Copyright Office)T1verified 2026.07.31
- Digital Replicas 101 (SAG-AFTRA)T1verified 2026.07.31
- Secret Invasion Opening Credits Generated By AI, Prompting Backlash (Deadline)T1verified 2026.07.31
- Brahma Announces the Acquisition of Metaphysic (DNEG)T1verified 2026.07.31
- MARZ Announces Breakthrough Vanity AI De-aging VFX System (Animation World Network)T2verified 2026.07.31
- Wonder Studio becomes Autodesk Flow Studio (CG Channel)T2verified 2026.07.31
- Beeble Studio Launches with Local 4K AI Relighting (CineD)T2verified 2026.07.31
- Topaz Video (Topaz Labs)T1verified 2026.07.31
- Flawless TrueSync (Flawless AI)T1verified 2026.07.31
- First AI Visually Dubbed Feature Film Hits US Theaters (New Atlas)T2verified 2026.07.31
- For the First Time, Netflix Uses GenAI for VFX in an Original Series (Interesting Engineering)T2verified 2026.07.31
- Late Night With the Devil Directors Clarify AI Image Use (Variety)T1verified 2026.07.31
- SAG-AFTRA Chief Lays Out What AI Protections It Wants In 2026 Contract (Deadline)T1verified 2026.07.31
Version 72
July 25, 2026
Upd
Update sweep and a deeper microdrama chapter
structuralA freshness sweep added Midjourney V8.2, Microsoft's MAI-Image-2.5-Pro and MAI-Voice-2-Flash preview, Meta's Muse Image and its consent controversy, and the Delhi High Court's ANI v. OpenAI fair-dealing ruling. The microdrama vertical got a deeper treatment: it monetizes like a free-to-play mobile game, its short vertical format fits exactly what AI video does well, AI arrived first through dubbing and localization, and distribution scale (ReelShort and DramaBox reportedly hold about seventy percent of revenue) is the moat, not the model. The ecosystem Microdramas directory was expanded accordingly.
Provenance
- Midjourney Version docs (V8 series)T1verified 2026.07.25
- Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash (Microsoft AI)T1verified 2026.07.25
- Meta rolls out Muse, a new AI image generator (TechCrunch)T1verified 2026.07.25
- ANI v. OpenAI: fair dealing and Indian copyright (SpicyIP)T2verified 2026.07.25
- Deloitte TMT Predictions 2026: Short-form video seriesT1verified 2026.07.25
- Could AI-Generated Microdramas Be the Future of Mobile Entertainment? (TheWrap)T2verified 2026.07.25
- Microdramas Go Global: the new wave of vertical video (Deadline)T2verified 2026.07.25
Version 68
July 25, 2026
Upd
Consistency fixes and a late-July currency pass
structuralA QA sweep reconciled internal contradictions introduced during the recent expansion (stale model hedges versus shipped notes, date and figure drift, Wan open-versus-closed, the Sora timeline in the Hollywood chapter). A freshness check against the leveled-up sources added dated notes: GPT Image 2 and Gemini Omni Flash now top the Artificial Analysis boards; the music industry launched a voluntary AI-Generated versus AI-Assisted labeling standard; Germany's GEMA v. Suno ruling is set for July 31, 2026; Udio added BuyDRM export enforcement; Netflix acquired Ben Affleck's InterPositive for about 587 million dollars; and Google DeepMind took a research stake in A24.
Ch. 69→Ch. 59→Ch. 56→Ch. 117→Ch. 46→
Provenance
- Artificial Analysis text-to-image leaderboardT2verified 2026.07.25
- Artificial Analysis text-to-video leaderboardT2verified 2026.07.25
- Music Community Introduces Labeling Program for Generative AI in Sound Recordings (RIAA)T1verified 2026.07.25
- GEMA, Suno copyright ruling set for July 31 (MLex)T1verified 2026.07.25
- Netflix paid $587M for Ben Affleck's AI filmmaking startup (Variety)T1verified 2026.07.25
- Google DeepMind and A24 announce research partnership (Google)T1verified 2026.07.25
- Udio fortifies its walled garden under BuyDRM deal (Digital Music News)T2verified 2026.07.25
- ByteDance Seed, Seedance 2.0 model pageT1verified 2026.07.25
Version 64
July 25, 2026
Upd
AI Filmmaking, roughly tripled and built on deep research
structuralExpanded the flagship AI Filmmaking part with new sections on consistency and control (image-first, reference conditioning, identity LoRAs, camera and keyframe control, performance transfer), sound and post tied to the audio pillar, where AI wins by format (ads, music video, the microdrama boom, previs, VFX), the studios and the money (Promise, Asteria and Moonvalley, Primordial Soup, Staircase, Toonstar, Critterz), festivals labor and the law (SAG-AFTRA consent, the Copyright Office report, the Thaler Supreme Court denial), and the real-time world-model frontier. Deepened the platform taxonomy and the verified filmmaker case studies.
Ch. 111→Ch. 113→Ch. 115→Ch. 117→Ch. 118→Ch. 120→
Provenance
- Kalshi airs AI-generated ad during NBA Finals using Google Veo 3 (Ad Age)T1verified 2026.07.25
- Runway Research: Introducing Runway Gen-4T1verified 2026.07.25
- Introducing Flow: Google's AI filmmaking tool designed for Veo (Google)T1verified 2026.07.25
- AI Studio Promise Raises Google, Ovitz's Crossbeam Investment (Hollywood Reporter)T1verified 2026.07.25
- Can Ethical AI Work in Hollywood? Moonvalley and Marey (TIME)T1verified 2026.07.25
- OpenAI-Backed Critterz Heads to Cannes Market (Deadline)T1verified 2026.07.25
- AI Film Critterz Looks for Tech Partner After OpenAI Sora shutdown (Bloomberg)T1verified 2026.07.25
- The Vertical Revolution: Global Microdrama Boom (Variety)T1verified 2026.07.25
- Copyright and Artificial Intelligence, Part 2: Copyrightability (US Copyright Office)T1verified 2026.07.25
- Supreme Court Denies Cert in AI Authorship Case (Mayer Brown)T2verified 2026.07.25
- WIRED: Behold Neural Viz, the first great cinematic universe of the AI eraT1verified 2026.07.25
- Runway releases its first world model, adds native audio (TechCrunch)T2verified 2026.07.25
Version 63
July 25, 2026
Upd
Editorial polish: audio naming and the frontier reframed as capstone
structuralClosed the editorial pass with structure and naming. The audio part is renamed Audio Models: Voice, Music, and Sound so its title reflects that audio is voice, music, and sound design, not voice alone. The audio-frontier section now sits with the other frontier topics. And the Research Frontier is reframed as the book's forward-looking capstone spanning image, video, and audio, resolving the jump from the industry chapters back into research methods.
Provenance
- SynthID (Google DeepMind)T1verified 2026.07.25
Version 62
July 25, 2026
Upd
Specialized Models deepened: instruction editing and lipsync
structuralDeepened two of the hottest specialized areas. Editing now covers the current instruction-based models (Gemini 2.5 Flash Image / Nano Banana, Qwen-Image-Edit, FLUX.1 Kontext, ByteDance SeedEdit and Seedream) and why surgical, multi-turn, consistent-character editing anchors real workflows. Lipsync was rebuilt from the Wav2Lip research lineage through Sync.so, Hedra Character-3, HeyGen, Runway Act-Two, and lip-sync now landing natively inside video models.
Provenance
- Introducing Gemini 2.5 Flash Image (Google Developers Blog)T1verified 2026.07.25
- Introducing FLUX.1 Kontext (Black Forest Labs)T1verified 2026.07.25
- ByteDance releases SeedEdit 3.0 (ByteDance Seed)T1verified 2026.07.25
- A Lip Sync Expert Is All You Need (arXiv, ACM MM 2020)T1verified 2026.07.25
- Creating with Act-Two (Runway Help Center)T1verified 2026.07.25
- Hedra Character-3: Omnimodal Character VideoT1verified 2026.07.25
Upd
Audio part deepened: native audio, Udio, and the music landscape
structuralMoved the Audio part closer to parity with image and video. Added a standalone section on native audio in video (Veo 3, Sora 2, Seedance, Wan, Kling, Grok) and the up-market pressure it creates. Deepened Udio (its fidelity edge, the October 2025 UMG settlement, and the download-lockout backlash) and the music and sound landscape (Lyria across Google's ecosystem, Google's acquisition of Riffusion via ProducerAI, Stable Audio's licensed-data play, AudioCraft's non-commercial weights, and Tencent and ACE-Step on the open frontier).
Provenance
- Google I/O 2025: new generative media models and toolsT1verified 2026.07.25
- Sora 2 is here (OpenAI)T1verified 2026.07.25
- UMG and Udio announce licensed AI music platform (PR Newswire)T1verified 2026.07.25
- Udio says users can download AI songs for 48 hours after backlash (Billboard)T2verified 2026.07.25
- Google acquires AI music platform ProducerAI, formerly Riffusion (Music Business Worldwide)T2verified 2026.07.25
- Stability AI introduces Stable Audio 2.5 (Stability AI)T1verified 2026.07.25
- AudioCraft: MusicGen, AudioGen, EnCodec (Meta AI)T1verified 2026.07.25
Version 61
July 25, 2026
Upd
AI Filmmaking rebuilt: from eleven stubs to six deep chapters
structuralConsolidated and deepened the flagship AI Filmmaking part. New arc: vibe directing as the new posture, the platforms and the agentic production loop, how an AI film actually gets made (with sound as a first-class stage tied to the audio pillar), the craft of technique-volume-and-the-edit, the films and filmmakers and where AI wins first, and the team of one and the author's chair. Verified examples include the Kalshi Veo 3 ad, Neural Viz, OpenArt's Director and vibe directing, Runway Gen-4/Aleph/Act-Two, the Runway AI Film Festival and Total Pixel Space, Critterz, the Here de-aging work, and Aronofsky's Primordial Soup.
Ch. 108→Ch. 109→Ch. 110→Ch. 112→Ch. 114→Ch. 119→
Provenance
- Kalshi airs AI-generated ad during NBA Finals using Google Veo 3 (Ad Age)T1verified 2026.07.25
- Vibe Directing: OpenArt launches Director (The Hollywood Reporter)T1verified 2026.07.25
- Behold Neural Viz, the first great cinematic universe of the AI era (WIRED)T1verified 2026.07.25
- Runway's AI Film Festival honors Total Pixel Space at Lincoln Center (Deadline)T1verified 2026.07.25
- Tom Hanks, Robin Wright de-aged in Zemeckis' Here using Metaphysic (The Hollywood Reporter)T1verified 2026.07.25
- Aronofsky's Primordial Soup debuts animated Revolutionary War series (Deadline)T1verified 2026.07.25
- OpenAI's animated film Critterz is coming (The Ankler)T2verified 2026.07.25
Version 60
July 25, 2026
New
Audio becomes first-class across the cross-cutting chapters
structuralEditorial pass stitching the audio pillar into the chapters that previously covered only image and video. Comparative Analysis now has an audio model landscape (voice latency, cloning, and the music structure and licensing axes). Evaluating Quality now has an audio quality section (naturalness, artifacts, consistency, controllability, intelligibility, provenance, weighted by use case). The Craft now covers directing voice, music, and sound, including the two-consents rule for cloning. The Operator's Playbook now models audio unit economics (speech meters on volume, music sells a cap), a hosted-vs-self-hosted decision, and the licensed-data-plus-indemnification moat. The Foundations gains a one-page audio-mechanics orientation.
Ch. 66→Ch. 87→Ch. 107→Ch. 09→Ch. 129→Ch. 142→
Provenance
- Moshi: a speech-text foundation model for real-time dialogue (arXiv)T1verified 2026.07.25
- F5-TTS: Fluent and Faithful Speech with Flow Matching (arXiv)T1verified 2026.07.25
- YuE: Scaling Open Foundation Models for Long-Form Music Generation (arXiv)T1verified 2026.07.25
- MMAudio: Multimodal Video-to-Audio Synthesis (arXiv)T1verified 2026.07.25
- Proactive Detection of Voice Cloning with Localized Watermarking / AudioSeal (arXiv)T1verified 2026.07.25
- Announcing Sonic: a low-latency voice model (Cartesia)T1verified 2026.07.25
- Models (ElevenLabs documentation)T1verified 2026.07.25
- Octave TTS (Hume blog)T1verified 2026.07.25
- Introducing Stable Audio 2.0 (Stability AI)T1verified 2026.07.25
- AudioCraft: MusicGen, AudioGen, EnCodec (Meta AI)T1verified 2026.07.25
- SynthID (Google DeepMind)T1verified 2026.07.25
- Voice cloning: how it works (ElevenLabs docs)T1verified 2026.07.25
- Warner Music Group and Suno Forge Partnership (PR Newswire)T1verified 2026.07.25
- Universal Music Group and Udio announce licensed platform (PR Newswire)T1verified 2026.07.25
- ElevenLabs Pricing (official)T1verified 2026.07.25
- Suno Pricing (official)T1verified 2026.07.25
Version 59
July 24, 2026
Upd
Figures for the audio foundations chapter
detailAdded two instrument-panel figures to the audio foundations chapter: the neural codec as the audio VAE (waveform to residual-vector-quantized tokens and back), and the two text-to-speech families (autoregressive codec language models versus flow matching).
Provenance
- High Fidelity Neural Audio Compression (EnCodec, arXiv 2210.13438)T1verified 2026.07.24
- F5-TTS: Flow Matching for Fluent and Faithful Speech (arXiv 2410.06885)T1verified 2026.07.24
Version 58
July 24, 2026
New
Audio part: the music landscape and the licensing reckoning
structuralFinal audio drop, completing the part. Added the music and sound landscape (Lyria, Stable Audio, Meta AudioCraft, Riffusion, Tencent, sound effects and foley, and the native-audio-in-video bifurcation that pushes standalone tools up-market) and a closing reckoning chapter on licensing and voice likeness: the turn toward licensed data, the ELVIS and NO FAKES laws and SAG-AFTRA consent regime, provenance via watermarking and C2PA, and the through-line that the ability to generate a voice does not settle whether it is lawful or right.
Provenance
- Stability AI introduces Stable Audio 2.5 (official)T1verified 2026.07.24
- AudioCraft: MusicGen, AudioGen, EnCodec (Meta AI)T1verified 2026.07.24
- Tennessee ELVIS Act analysis (Skadden)T1verified 2026.07.24
- SAG-AFTRA Artificial Intelligence resources (official)T1verified 2026.07.24
Version 57
July 24, 2026
New
Audio part: the voice landscape, Suno, and Udio
structuralSecond audio drop. Added the voice landscape beyond ElevenLabs (the real-time lane with Cartesia, OpenAI, Google, Hume, Sesame and the Scarlett Johansson Sky episode; the open lane with Kokoro, Chatterbox, Fish Speech and Microsoft declining to release VALL-E 2; the editing and enterprise lane; and voice in filmmaking, where speech-to-speech conversion beats text-to-speech), plus full deep dives on Suno and Udio, including their mirror-image settlement paths with the record labels.
Provenance
- Cartesia Sonic (official)T1verified 2026.07.24
- Scarlett Johansson responds to OpenAI Sky voice (Variety)T1verified 2026.07.24
- Record Companies Bring Cases Against Suno and Udio (RIAA)T1verified 2026.07.24
- UMG and Udio announce licensed AI music platform (PR Newswire)T1verified 2026.07.24
Version 56
July 24, 2026
New
Audio and Voice becomes a first-class part (foundations + ElevenLabs)
structuralAdded a new Audio and Voice part alongside image and video, opening with a foundations chapter on how audio generation works (neural codecs as the audio VAE, the two TTS families and voice cloning, why music is harder than speech, sound effects, and native audio in video) and a full ElevenLabs deep dive covering the company, the model line and product surface, getting the best output and holding a character voice across a film, and the likeness and consent problem. Voice, music, and the licensing reckoning follow in subsequent drops.
Provenance
- High Fidelity Neural Audio Compression (EnCodec, arXiv 2210.13438)T1verified 2026.07.24
- VALL-E: Neural Codec Language Models are Zero-Shot TTS (arXiv 2301.02111)T1verified 2026.07.24
- ElevenLabs raises $500M Series D at $11B valuation (official)T1verified 2026.07.24
- Eleven v3: Most Expressive AI TTS Model (official)T1verified 2026.07.24
Version 55
July 24, 2026
Upd
Training-data FAQ: the local versus hosted distinction
structuralAnswered the missing case in the will-my-prompts-train-the-next-model FAQ: hosted company UIs and APIs run on the providers servers and can log and train on your inputs, while running open-weight models locally through something like ComfyUI keeps everything on your machine, private by construction rather than by promise.
Provenance
- Midjourney Terms of ServiceT1verified 2026.07.24
- ComfyUI (official GitHub repository)T1verified 2026.07.24
Version 54
July 24, 2026
Upd
ComfyUI walkthrough, now with a worked example
structuralRewrote the ComfyUI workflow walkthrough to follow one concrete image (a cinematic astronaut in a wheat field, made with an SDXL checkpoint like Juggernaut XL) through every node, with specific prompts, a 25-step DPM++ 2M Karras KSampler setup, and named variations (image-to-image, LoRA, upscale, ControlNet, and swapping in a Wan or Hunyuan video loader), so the abstract graph is grounded in an actual use case.
Provenance
- ComfyUI (official GitHub repository)T1verified 2026.07.24
- ComfyUI Official DocumentationT1verified 2026.07.24
Version 53
July 24, 2026
Upd
ComfyUI, for beginners: why it feels so complicated
structuralAdded a subsection to the ComfyUI infrastructure chapter that meets a beginner's overwhelm head on, with analogies (the airplane cockpit, the microwave versus the professional kitchen, the graph as a language) and the core why: ComfyUI does not add complexity, it reveals the complexity that was always there, and every workflow is the same six-node skeleton with small insertions.
Provenance
- ComfyUI (official GitHub repository)T1verified 2026.07.24
- ComfyUI Official DocumentationT1verified 2026.07.24
Version 50
July 24, 2026
Upd
Krea, in depth: aggregator, canvas, and model lab
detailDeepened the Krea chapter with its honest dual nature as a real-time canvas and aggregation layer that also trains its own models, the flag that Krea 1 rode the FLUX ecosystem and Krea Realtime 14B distilled Wan while only Krea 2 is from scratch, its real-time and aesthetic-post-training edge, and practical access and workflow guidance.
Provenance
- Investing in Krea (Andreessen Horowitz)T1verified 2026.07.24
- Releasing Open Weights for FLUX.1 Krea (Krea blog)T1verified 2026.07.24
- Krea 2 Raw and Turbo available as open weights (VentureBeat)T1verified 2026.07.24
Upd
Qwen-Image, in depth: the open bet and the closed drift
detailDeepened the Alibaba Qwen-Image chapter with the 20B MMDiT on a frozen Qwen2.5-VL encoder, the Apache 2.0 open weights and standout Chinese text rendering, the honest 2026 drift as Qwen-Image-3.0 shipped closed with no weights or report, and practical open-weight, quantized-local, and API access guidance.
Provenance
- Qwen-Image Technical Report (arXiv 2508.02324)T1verified 2026.07.24
- Qwen/Qwen-Image model card (Hugging Face)T1verified 2026.07.24
- Alibaba Launches Qwen-Image-3.0 Without Benchmarks or Weights (Unite.AI)T1verified 2026.07.24
Upd
Seedream, in depth: the unified 4K engine
detailDeepened the ByteDance Seedream chapter with the unified Seedream 4.0 (one diffusion transformer for generation, editing, and composition at 4K in about 1.8 seconds), the launch arena double-top over Nano Banana that later eroded, the secondary-sourced 4.5 and 5.0 flags, and practical access across Dreamina, Doubao, and BytePlus with the US caveat.
Provenance
- Seedream 4.0: Toward Next-generation Multimodal Image Generation (arXiv 2509.20427)T1verified 2026.07.24
- Seedream 4.0 Officially Released (ByteDance Seed blog)T1verified 2026.07.24
- Artificial Analysis Text-to-Image LeaderboardT2verified 2026.07.24
Upd
GPT-Image, in depth: the native turn and the Ghibli moment
detailDeepened the OpenAI image chapter with the March 2025 shift from standalone DALL-E to native generation inside GPT-4o, the viral Studio Ghibli moment and the melting-GPUs demand, the gpt-image-1 through gpt-image-2 lineage now topping the arena, and practical access, token pricing, and the honest weaknesses (speed, the warm tint, moderation).
Provenance
- Introducing 4o Image Generation (OpenAI)T1verified 2026.07.24
- Sam Altman says ChatGPT's Ghibli-style images are melting OpenAI's GPUs (Fortune)T1verified 2026.07.24
- Artificial Analysis Text-to-Image LeaderboardT2verified 2026.07.24
Version 49
July 24, 2026
Upd
Stability, in depth: the collapse, the license, and the reset
detailExpanded the thin Stability chapter into a full profile: the SD3 license backlash that drove the community to FLUX, the two architectural eras (latent diffusion to MMDiT rectified flow), honest strengths and weaknesses, the near-bankruptcy and the Akkaraju, Parker, and Cameron reset, practical open-weight and licensing guidance, and the honest flag that the Stable Diffusion 4 reports are SEO blog spam.
Provenance
- How Stability AI's Founder Tanked His Billion-Dollar Startup (Forbes)T1verified 2026.07.24
- Introducing Stable Diffusion 3.5 (Stability AI)T1verified 2026.07.24
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (arXiv 2403.03206)T1verified 2026.07.24
Upd
Midjourney, in depth: the lawsuits and how to use it
detailDeepened the Midjourney chapter with the Disney, Universal, and Warner Bros. copyright lawsuits and the training-data opacity behind them, practical access and parameter prompting guidance (--ar, --stylize, --sref, RAW), and the bootstrapped, profitable independence funding a long arc toward video, 3D, and world simulation.
Provenance
- Disney and Universal sue Midjourney for AI copyright infringement (CNBC)T1verified 2026.07.24
- Warner Bros. Discovery Sues Midjourney (Deadline)T1verified 2026.07.24
- Midjourney launches its first AI video generation model, V1 (TechCrunch)T1verified 2026.07.24
Upd
Recraft, in depth: the design vertical
detailDeepened the Recraft chapter with its design-and-brand positioning, the red panda stealth arena win with V3, the editable-vector and typography differentiators, the V4 and V4.1 design-taste rebuilds, and practical guidance on when to choose it over general image models.
Provenance
- A stealth AI model beat DALL-E and Midjourney, its creator landed 30M (TechCrunch)T1verified 2026.07.24
- Introducing Recraft V4: Design Taste Meets Image Generation (Recraft blog)T1verified 2026.07.24
- Recraft V3 model page (Artificial Analysis)T2verified 2026.07.24
Upd
Ideogram, in depth: the eroding typography moat
detailDeepened the Ideogram chapter with the honest read that its text-rendering moat narrowed as DALL-E, Imagen, FLUX, Recraft, and GPT-native generation caught up, the gap between its own and neutral benchmark rankings, the unverified Ideogram 4.0 open-weight reports, and practical prompting for legible in-image text.
Provenance
- Midjourney rival Ideogram gets 80M Series A led by a16z (VentureBeat)T1verified 2026.07.24
- Ideogram 3.0 gets a realism boost and new editing tools (The Decoder)T1verified 2026.07.24
- Artificial Analysis Text-to-Image LeaderboardT2verified 2026.07.24
Upd
Nano Banana, in depth: the Gemini-native consolidation
detailDeepened the Google image chapter with the consolidation onto the Gemini-native path (Imagen 4 retirement, migration to Nano Banana 2), the confusing Pro-versus-2 tier naming, the mid-2026 arena standing behind GPT Image 2, and practical guidance on the Pro, 2, and 2 Lite ladder plus the identity-consistency and conversational-editing strengths.
Provenance
- Introducing Gemini 2.5 Flash Image (Google Developers Blog)T1verified 2026.07.24
- How Nano Banana got its name (Google blog)T1verified 2026.07.24
- Artificial Analysis Text-to-Image LeaderboardT2verified 2026.07.24
Upd
FLUX, in depth: the multimodal turn and how to use it
detailDeepened the Black Forest Labs FLUX chapter with the 2026 multimodal turn (FLUX.2, the July FLUX 3 unified image, video, audio, and action model, the FLUX mimic robotics model and Audi test, and the 300-million-dollar Series B) plus practical open-versus-pro access and licensing guidance.
Provenance
- FLUX 3: Towards Multimodal Flow Models as the Backbone of Visual Intelligence (BFL blog)T1verified 2026.07.24
- Black Forest Labs raises 300M at 3.25B valuation (TechCrunch)T1verified 2026.07.24
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing (arXiv 2506.15742)T1verified 2026.07.24
Version 48
July 24, 2026
Upd
Grok Imagine, in depth: access and the moderation problem
detailAdded practical access, pricing, and prompting guidance to the xAI Grok Imagine chapter, plus an honest, sourced account of the permissive-moderation stance and the nonconsensual-deepfake controversy and regulatory response that followed.
Provenance
- Introducing Grok Imagine 1.0 (xAI official post)T1verified 2026.07.24
- Musk's AI chatbot Grok is still making sexual deepfakes (NBC News)T1verified 2026.07.24
- xAI updates Grok Imagine to 1.5 with image to video at 720p (the-decoder)T2verified 2026.07.24
Upd
Pika, in depth: the consumer bet
detailDeepened the Pika chapter with the Guo and Meng founding and their diffusion-research pedigree, the deliberate choice to own the fun, social, effects-driven corner of AI video rather than the frontier, honest strengths and weaknesses, and practical guidance on using the effects and ingredient features.
Provenance
- Pika Labs raises 55M (TechCrunch)T1verified 2026.07.24
- Pika 2.0 launches in wake of Sora (VentureBeat)T1verified 2026.07.24
- Pika, a TikTok-like AI app built for Gen Z (Fortune)T1verified 2026.07.24
Upd
Luma, in depth: the world-model turn and the HUMAIN raise
detailDeepened the Luma chapter with its pivot from video generation to unified world models (Ray3 reasoning, Uni-1 and Luma Agents), the 900-million-dollar HUMAIN Series C and the Project Halo supercluster, honest strengths and weaknesses, and practical access and HDR-pipeline guidance.
Provenance
- Luma AI launches Ray3 (official)T1verified 2026.07.24
- Luma AI Raises 900 Million Series C Led by HUMAIN (PRNewswire)T1verified 2026.07.24
- Luma launches creative AI agents powered by Unified Intelligence (TechCrunch)T1verified 2026.07.24
Upd
Wan, in depth: the open-core split
detailExpanded the Alibaba Wan chapter with the crucial 2026 finding that its open weights stop at Wan 2.2 while 2.5, 2.6, and likely 2.7 are closed API-only products (an open-core split, not full openness), plus honest strengths and weaknesses and open-versus-hosted access guidance.
Provenance
- Wan: Open and Advanced Large-Scale Video Generative Models (arXiv 2503.20314)T1verified 2026.07.24
- Wan-AI organization on Hugging Face (official open weights)T1verified 2026.07.24
- Alibaba Unveils Wan2.6 Series (Alibaba Cloud)T1verified 2026.07.24
Upd
Hailuo, in depth: the honest read and the studio lawsuit
detailAdded an honest strengths-and-weaknesses read to the MiniMax Hailuo chapter (the stale mid-2025 arena ranking, the 456B parameter figure that is actually the M1 language model, and the Disney, Universal, and Warner lawsuit) plus practical access and director-camera prompting guidance.
Provenance
- MiniMax Hailuo 02 (MiniMax official news)T1verified 2026.07.24
- Disney, Universal, Warner Bros. Discovery sue China's MiniMax (CNBC)T1verified 2026.07.24
- Artificial Analysis Text to Video LeaderboardT2verified 2026.07.24
Upd
Seedance, in depth: the four-name model and how to use it
detailAdded practical access guidance (Dreamina, Jimeng, Doubao, CapCut, BytePlus, and the US-availability caveat), the honest resolution-spec inconsistency, and the unified-multimodal direction including Seedance 2.0's native audio, to the ByteDance Seedance chapter.
Provenance
- Seedance 1.0: Exploring the Boundaries of Video Generation Models (arXiv 2506.09113)T1verified 2026.07.24
- ByteDance Seed Seedance product pageT1verified 2026.07.24
- ByteDance's Dreamina Seedance 2.0 comes to CapCut (TechCrunch)T1verified 2026.07.24
Version 46
July 24, 2026
Upd
The ComfyUI field manual, made understandable
detailRewrote the densest ComfyUI passages for a general reader. Each hard section now opens with an everyday analogy before the machinery: the installed tool as a small power plant you run rather than a storefront you visit, the typed wires as plumbing that only connects when the gauges match, the six-node default workflow as a darkroom assembly line where the painter works on a compressed draft (which is why latents, samplers, and a VAE decode each exist), the denoise dial as how much of the canvas you repaint, the API export as turning a hand-drawn recipe into a callable button, and owning versus renting a GPU as owning versus renting a car. New figures render the six typed data wires, the model-to-workflow-to-agent climb, and the four ways to run, and the canonical graph and latent-pipeline diagrams now sit inside the manual itself.
Ch. 70→Ch. 71→Ch. 72→Ch. 73→Ch. 77→Ch. 79→
Provenance
- ComfyUI documentation (official)T1verified 2026.07.24
- ComfyUI (GitHub)T1verified 2026.07.24
Upd
FLUX 3: Black Forest Labs goes multimodal
structuralNoted the July 23 launch of FLUX 3, a unified image, video, audio, and action model with native synchronized audio and a companion robotics model (FLUX mimic), in the Black Forest Labs chapter. This is the first finding promoted out of the weekly Field Reports into the standing book.
Provenance
- Black Forest Labs Unveils FLUX 3, A New Multimodal Frontier ModelT1verified 2026.07.24
- Black Forest Labs launches FLUX 3 (images and 20-second video with audio)T1verified 2026.07.24
- Black Forest Labs Unveils First Model for Robotics in Shift to Physical AI (Bloomberg)T1verified 2026.07.24
Version 45
July 23, 2026
New
Higgsfield, in depth: the aggregator bet and its honest read
detailAdded a Higgsfield profile to the ecosystem aggregator layer: the cinematic camera-control differentiator, the DoP and Soul models wrapping frontier engines like Kling, Veo, and Sora, the steep 2026 revenue and funding ramp toward a reported 1.3 billion valuation, and an honest read on the wrapped-model quality ceiling and the February 2026 Forbes controversy over fake demos, deepfake content, and billing.
Provenance
- Higgsfield coverage (funding, $1.3B valuation) - TechCrunchT1verified 2026.07.23
- Racist Videos And Payment Problems: The Dark Side Of This AI Startup's Super-Fast Growth - ForbesT1verified 2026.07.23
- This AI video startup is having its breakout moment ($500M revenue) - Business InsiderT1verified 2026.07.23
- Higgsfield AI (official site) products and modelsT1verified 2026.07.23
Version 44
July 23, 2026
Upd
Runway, in depth: the origin, the moat, and how to use it
detailDeepened the Runway chapter with the company's NYU ITP founding and its co-authorship of Stable Diffusion, the funding arc and the AI Film Festival, an honest read on where its editing-first control beats raw generation and where it trails the frontier, and a practical guide to plans, pricing, and treating Runway as an editing and orchestration hub rather than a clip slot machine.
Provenance
- Introducing Runway Aleph (official)T1verified 2026.07.23
- Runway Pricing (official)T1verified 2026.07.23
- Runway (company) - WikipediaT2verified 2026.07.23
Version 43
July 23, 2026
Upd
Veo, in depth: the org, the moat, and getting the best out of it
detailDeepened the Google Veo chapter with the DeepMind organization and the vertical-integration moat (in-house TPUs, the YouTube corpus, and distribution across Flow, Gemini, Vertex AI, and the API), the reported seventy-five million dollar A24 partnership, and a practical guide to the four access paths, per-second and credit pricing, and how to prompt for camera, lens, and native audio.
Provenance
- Veo - Google DeepMind model pageT1verified 2026.07.23
- Gemini API pricing (Veo and Gemini Omni Flash video rates)T1verified 2026.07.23
- Veo (text-to-video model) - WikipediaT2verified 2026.07.23
- A24 - Wikipedia (Google investment, June 2026)T2verified 2026.07.23
Version 42
July 23, 2026
Upd
Sora, in depth: the team, the world-simulator thesis, and the shutdown
detailRebuilt the OpenAI Sora chapter into a full profile: Tim Brooks and Bill Peebles and the DiT paper Sora was built on, the spacetime-patch world-simulator thesis read honestly against what actually emerged, Sora 2's native audio and Cameos, and the economics and cultural legacy behind the 2026 app shutdown and API sunset.
Provenance
- Video generation models as world simulators (OpenAI Sora technical report)T1verified 2026.07.23
- Scalable Diffusion Models with Transformers (DiT, Peebles and Xie, arXiv 2212.09748)T1verified 2026.07.23
- Sora (text-to-video model) - WikipediaT2verified 2026.07.23
Version 41
July 23, 2026
Upd
Kling, in depth: the business, the tech, and where it stands
detailExpanded the Kling chapter with Kuaishou's business and the July 2026 spin-off (a ~$2.8B raise near an $18B valuation, Kuaishou keeping ~68%, Tencent/Alibaba/Baidu backing), an honest architecture read from the Kling-Omni and Kling-Foley papers, its mid-2026 arena standing (~#6, behind Gemini Omni Flash, Seedance, and Wan), and practical guidance on modes, access, pricing, and prompting.
Provenance
- Kuaishou restructures Kling AI (spin-off, valuation, financials)T2verified 2026.07.23
- Kling AI ARR hits USD240M in December 2025 (Kuaishou official)T1verified 2026.07.23
- Kling-Omni Technical Report (arXiv 2512.16776)T1verified 2026.07.23
- Artificial Analysis Video ArenaT2verified 2026.07.23
Version 40
July 23, 2026
Upd
Driving ComfyUI agentically with MCP
detailComfy Org shipped an official hosted Comfy Cloud MCP server (public beta, June 2026) that lets an AI assistant search, assemble, run, and iterate on ComfyUI workflows, alongside a growing set of local community MCP servers (Artokun's control plane, Joe Norton's lightweight server) that can also edit the live graph and install nodes and models.
Provenance
- Comfy MCP: Turn your agent into a creative technologistT1verified 2026.07.23
- Comfy Cloud MCP, official ComfyUI docsT1verified 2026.07.23
- artokun/comfyui-mcp (local agent control plane)T1verified 2026.07.23
Version 32
July 23, 2026
Upd
ComfyUI's 2026 releases: quantization, SeedVR2, and new conditioning
detailComfyUI added int4/int8 quantization, native SeedVR2 upscaling, Depth Anything 3, SCAIL-2 multi-reference character replacement, and PixelDiT support through mid-2026, while ai-toolkit added LoRA training for FLUX.2 and Z-Image.
Provenance
- ComfyUI ChangelogT1verified 2026.07.23
- ostris/ai-toolkitT1verified 2026.07.23
Version 26
July 23, 2026
Upd
AI copyright suits advance; EU AI Act enforcement begins August 2
detailA judge let the studios' suit against MiniMax's Hailuo proceed (May 26), the New York Times sought sanctions against OpenAI over discovery (July), and the EU AI Act's enforcement powers over general-purpose model providers take effect August 2, 2026 with fines up to 15M euros or 3% of global turnover.
Provenance
- Midjourney/MiniMax studio copyright litigation (Variety)T2verified 2026.07.23
- NYT-led group asks court to sanction OpenAI (Reuters)T1verified 2026.07.23
- Enforcement of Chapter V under the EU AI ActT2verified 2026.07.23
Upd
Studios move from lawsuits to equity: A24, Lionsgate, Bertelsmann, Getty
detailMid-2026 saw rights-holders take stakes and sign licensing deals: Google DeepMind invested $75M in A24 to co-develop Veo tools, Lionsgate took an equity stake in Runway (with a Bertelsmann partnership following), and Getty Images signed a display-only deal with OpenAI for ChatGPT.
Provenance
- A24 and Google DeepMind form AI PartnershipT1verified 2026.07.23
- Lionsgate Takes Equity Stake in Runway AIT1verified 2026.07.23
- Getty Images Announces Display Partnership with OpenAIT1verified 2026.07.23
Version 24
July 23, 2026
Upd
Google's Nano Banana 2 rolls out in three tiers
detailGoogle shipped Nano Banana 2 (Gemini 3.1 Flash Image) in February 2026 across Gemini, Search, and Ads, in Lite, standard, and Pro tiers, with several variants ranking in the image-arena top ten.
Provenance
- Google Blog: Nano Banana 2T1verified 2026.07.23
Upd
Ideogram 4.0 ships as the lab's first open-weight model
detailIdeogram released Ideogram 4.0 on June 3, 2026: a 9.3B single-stream diffusion Transformer with a Qwen3-VL text encoder, native 2K resolution, and JSON prompting, its first model with open weights alongside API access.
Provenance
- Ideogram 4.0T1verified 2026.07.23
Upd
Black Forest Labs launches FLUX 3, its first multimodal model
detailFLUX 3 launched July 23, 2026 as a multimodal frontier model generating images plus video with synchronized audio (clips up to roughly 20 seconds), in limited early access, extending Black Forest Labs beyond still images.
Provenance
- Black Forest Labs BlogT1verified 2026.07.23
- Black Forest Labs launches FLUX 3 (VentureBeat)T2verified 2026.07.23
Upd
GPT Image 2 tops the image arena
detailOpenAI released GPT Image 2 (ChatGPT Images 2.0) on April 21, 2026, retiring DALL-E 3 and GPT-Image-1.5 with near-perfect text rendering, higher resolution, and a thinking mode. It currently ranks first on the Artificial Analysis text-to-image arena.
Provenance
- Introducing ChatGPT Images 2.0T1verified 2026.07.23
- Artificial Analysis Text-to-Image LeaderboardT2verified 2026.07.23
Version 23
July 23, 2026
Upd
Real-time world models draw funding and open releases
detailWorld models became the fastest-moving frontier category in 2026: Reactor emerged from stealth with a $59M round to build a developer platform for real-time AI worlds, alongside World Labs' Marble, Google DeepMind's Genie 3, and Tencent's open real-time HY-World.
Provenance
- Real-Time AI Video Startup Reactor Raises $59 MillionT1verified 2026.07.23
Upd
The inference layer's 2026 funding surge
detailThe compute layer that serves generative media drew huge rounds in mid-2026: fal named AWS its preferred cloud at roughly 2.5 million developers, Baseten raised a $1.5B Series F near a $13B valuation, and Together AI raised $800M at $8.3B, all on rapidly growing inference demand.
Provenance
- fal Scales the World's Largest Generative Media Platform with AWST1verified 2026.07.23
- Announcing our Series F (Baseten)T1verified 2026.07.23
- Together AI raises $800M, leaps to $8.3B valuationT1verified 2026.07.23
Version 22
July 23, 2026
Upd
Kuaishou weighs a Kling spin-off near a $20B valuation
detailKuaishou disclosed in a May 12 exchange filing that it is evaluating a restructuring of the Kling AI business, with reports pointing to a pre-IPO round near a twenty-billion-dollar valuation as Kling's revenue run rate climbed toward five hundred million dollars.
Provenance
- Kuaishou Weighs Seeking Outside Money for AI Video UnitT2verified 2026.07.23
Upd
Luma's Ray3.2 adds frame-level control and EXR output
detailLuma released Ray3.2 (June 9), adding frame-by-frame directability and HDR generation with paired EXR outputs for professional pipelines, followed by 'Skills' repeatable agent workflows (June 16).
Provenance
- Introducing Ray3.2T1verified 2026.07.23
Upd
Seedance 2.5 claims a thirty-second single take
detailByteDance unveiled Seedance 2.5 at its Volcano Engine FORCE conference (June 23), claiming native single-shot video up to thirty seconds and up to roughly fifty reference inputs, with public API access via BytePlus opening July 16.
Provenance
- ByteDance's Seedance 2.5 breaks the 30-second barrierT2verified 2026.07.23
Upd
Runway's 2026 turn: Aleph 2.0, a Media Router, and studio deals
detailRunway shipped Aleph 2.0 and Edit Studio (May 21), then a Media Router that selects across image, video, and audio models (July 23), while taking a Lionsgate equity stake (June 11) and announcing a Bertelsmann partnership (July 1). The pattern is a bet on owning workflow and distribution rather than only the base model.
Provenance
- Introducing Aleph 2.0 and Edit StudioT1verified 2026.07.23
- Runway bets on AI model routing as generative media gets crowdedT2verified 2026.07.23
- Lionsgate Takes Equity Stake in Runway AIT1verified 2026.07.23
Upd
Google moves video into Gemini Omni; Omni Flash tops the arenas
detailGoogle shipped no Veo 4 at I/O 2026 and instead launched Gemini Omni Flash (June 30), a fast video generation and conversational-editing model that now ranks first on both the text-to-video and image-to-video arenas. Veo 3.1 remains the standalone flagship.
Provenance
- Watch 9 Google videos of Gemini Omni and Gemini 3.5T1verified 2026.07.23
- Artificial Analysis Video ArenaT2verified 2026.07.23
Version 21
July 23, 2026
New
ComfyUI, from setup to senior
structuralThe ComfyUI manual gains a mastery arc: where it runs (local vs rented GPUs, RunPod, serverless, Comfy Cloud), the pitfalls that cost beginners a weekend, the 80/20 path to getting good, how professionals actually build with it, and what 'senior at ComfyUI' means.
Ch. 79→Ch. 80→Ch. 81→Ch. 82→Ch. 83→
Provenance
- ComfyUI documentationT1verified 2026.07.23“Official node reference, troubleshooting, custom-node and API docs.”
- Comfy Org, subgraphs official releaseT1verified 2026.07.23“Package related nodes into a single clean subgraph node, turning them into LEGO blocks.”
- RunPod, run ComfyUI on a PodT1verified 2026.07.23“Launch a ComfyUI template and connect to the HTTP service on port 8188.”
New
In the studio: how the best AI creators actually work
structuralTwo chapters on process: deep profiles of how PJ Accetturo, Neural Viz, Dave Clark, Paul Trillo, and the Dor Brothers actually make their work, and the patterns that recur across all of them (the common stack, generate-then-cull ratios, model-mixing, and editing discipline).
Provenance
- Neural Viz (channel)T1verified 2026.07.23“The Monoverse: an ongoing AI-made alien television universe, produced largely by one person.”
- WIRED, on Neural Viz (Beam)T3verified 2026.07.23“A longtime filmmaker who voices every character himself, using Midjourney, ElevenLabs, and Runway.”
- Business Insider, on PJ Accetturo's Kalshi adT3verified 2026.07.23“300 to 400 generations to get 15 usable clips; I always tell it to return 5 prompts at a time.”
- Every, how a Hollywood director uses AI (Shipper/Clark)T3verified 2026.07.23“You gotta do a bunch of generations; take parts from each clip.”
- Forbes, Charlie Fink on Paul TrilloT3verified 2026.07.23“Roughly 700 clips generated, 55 kept, for a continuous infinite-dolly music video.”
Version 20
July 22, 2026
New
Inside fal: a tour of the model garden
structuralA new chapter uses fal, the developer inference platform hosting 1,000+ generative models, as a vantage point on the whole medium: a guided tour of the interesting and less-mainstream models by category, profiling Bria (licensed-data image), Reve, Sonilo (video-to-music), Decart (real-time world models), and Lightricks' open LTX-2, among many others.
Provenance
- fal Series C announcementT1verified 2026.07.22“The Generative Media Platform for Developers; in the last 12 months our revenue grew 60x.”
- fal (models and platform)T1verified 2026.07.22“The world's best generative image, video, and audio models, all in one place; 1,000+ production-ready models.”
- Bria (licensed-data visual AI)T1verified 2026.07.22“Enterprise visual AI trained on 100% licensed data with full IP indemnification and an attribution engine.”
- Decart publications (real-time world models)T1verified 2026.07.22“The infrastructure and models that make AI run at the speed of reality; Oasis and MirageLSD.”
- Lightricks LTX-2T1verified 2026.07.22“Open-weights synchronized audio and video at native 4K, runnable on consumer GPUs.”
Version 19
July 22, 2026
Upd
Freshness pass: current model versions and arena standings
detailVerified the mid-2026 landscape against the live leaderboards and updated the stragglers: Midjourney V8.1 is now the default, ByteDance shipped Seedream 5.0 Pro, and the current arena order (GPT Image 2 on top, Reve 2.1 and Microsoft MAI-Image-2.5 behind; Gemini Omni Flash leading video) is confirmed.
Provenance
- Artificial Analysis text-to-image arenaT2verified 2026.07.22“GPT Image 2 leads the arena, with Reve 2.1 and Microsoft MAI-Image-2.5 clustered just behind.”
- Artificial Analysis text-to-video arenaT2verified 2026.07.22“Google Gemini Omni Flash leads text-to-video and image-to-video.”
- Midjourney version historyT1verified 2026.07.22“V8.1 released and became the default model in mid-2026.”
Version 18
July 22, 2026
New
A ComfyUI field manual
structuralA thorough hands-on manual for ComfyUI in The Machine: install and setup, reading the node graph, the default text-to-image workflow node by node, the core recipes (img2img, inpaint, outpaint, upscale, LoRA, ControlNet, IPAdapter), native video and audio, the Manager and registry, the headless/agentic API, and practical memory and reproducibility.
Ch. 70→Ch. 72→Ch. 73→Ch. 74→Ch. 75→Ch. 77→Ch. 78→
Provenance
- ComfyUI (GitHub)T1verified 2026.07.22“The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.”
- ComfyUI documentationT1verified 2026.07.22“Official node reference and tutorials, including the text-to-image graph and KSampler parameters.”
- Comfy Org raises $17MT1verified 2026.07.22“Open source must win. If a proprietary service dominates, creativity loses.”
New
AI filmmaking, deeper: platform mechanics, where it wins, and the team of one
structuralThree more chapters on the craft: how the agentic loop actually runs inside Director, Flow, LTX Studio, and Showrunner; where AI filmmaking is already winning (ads, music videos, microdramas, previs, VFX) and where it still loses; and the real team shape and economics of an AI production, from Critterz to the solo creator.
Provenance
- Google, Introducing FlowT1verified 2026.07.22“Flow is a new AI filmmaking tool built with creatives for the next wave of storytelling.”
- The Hollywood Reporter, OpenArt's Director (Zeitchik)T3verified 2026.07.22“Users can revise films without full regeneration; directing it is like giving notes to an editor.”
- Variety, Lionsgate takes equity stake in RunwayT3verified 2026.07.22“Lionsgate expands its Runway partnership from previz and post into an equity stake.”
- The Hollywood Reporter, Aronofsky's AI series (Han)T3verified 2026.07.22“A flat, plasticky sheen; a hairbrush that glides along a lock of hair without actually going through it.”
- TheWrap, AI-generated microdramas (Patton)T3verified 2026.07.22“AI-made vertical series targeting roughly $10,000 per hour of content versus about $150,000 for live-action.”
Version 17
July 22, 2026
New
The book, restructured into two movements, and a new Ecosystem chapter
structuralThe book is now organized as two movements, The Machine (how generative models work) and The Medium (what is being built on them). A new chapter walks the entire creative stack in prose, phase by phase, naming the companies doing the work in each box.
Ch. 95→Ch. 96→Ch. 97→Ch. 98→Ch. 99→Ch. 100→
Provenance
- Machine Cinema (market map)T2verified 2026.07.22“A market map of the AI creative ecosystem across development, production, post, inference, and distribution.”
New
AI filmmaking, in depth: the pipeline, the techniques, and the films
structuralThe AI Filmmaking part gains three craft chapters: how an AI film actually gets made end to end, the techniques that separate finished work from slop, and the real films and filmmakers driving the moment (PJ Accetturo, Neural Viz, Curious Refuge, Runway's festival, and the Sora-made feature Critterz).
Provenance
- OpenArt, What Is Vibe Directing?T1verified 2026.07.22“Vibe directing is directing a video into existence by talking. You shape every shot and cut with your feedback in chat.”
- Google, Flow: an AI filmmaking toolT1verified 2026.07.22“Flow is a new AI filmmaking tool built with creatives for the next wave of storytelling.”
- Runway AI Film Festival 2025T1verified 2026.07.22“Third edition; thousands of submissions; finalists screened at IMAX; Grand Prix to Total Pixel Space.”
- The Hollywood Reporter, on Curious Refuge (Hibberd)T3verified 2026.07.22“The biggest misconception about AI filmmaking is you type in a prompt and get a film. It's artistry.”
- WIRED, on Neural Viz (Beam)T3verified 2026.07.22“A longtime filmmaker who voices all the characters himself, using Midjourney, ElevenLabs, and Runway.”
Version 16
July 22, 2026
New
The creator scene, the memes, and the slop fight
structuralThe native AI-creator canon (Will Smith spaghetti, the Ghibli trend, Neural Viz, PJ Ace, Curious Refuge, Runway's AIFF) and the two backlash fights, over slop (Willison's definition) and over consent (Glaze/Nightshade, Tilly Norwood, ScarJo vs Sky, likeness vaults, the jobs debate).
Provenance
- Simon Willison: SlopT1verified 2026.07.22“'Slop is the new name for unwanted AI-generated content'; sharing unreviewed AI content is rude.”
- SAG-AFTRA on Tilly Norwood (Variety)T3verified 2026.07.22“SAG-AFTRA: Tilly Norwood 'is not an actor... trained on the work of countless professional performers without permission or compensation.'”
- Will Smith Eating Spaghetti test (Wikipedia)T2verified 2026.07.22“The community's running benchmark for AI video progress since March 2023.”
New
The deals: Lionsgate, Netflix, Disney-OpenAI, the Sphere
structuralWhat the studios actually did: Lionsgate-Runway's underdelivery and restructuring, Netflix's ~300 AI-assisted 2026 titles, the Disney-OpenAI Sora deal and its March 2026 collapse plus Disney's lawsuits, and the Sphere's $370M+ AI Wizard of Oz.
Provenance
- Runway and Lionsgate expand partnershipT1verified 2026.07.22“The 2024 custom-model deal restructured in June 2026 into an equity stake and co-production.”
- Netflix ~300 AI titles in 2026 (Variety)T3verified 2026.07.22“Netflix: about 300 of its 2026 titles used generative-AI-assisted workflows.”
- Disney-OpenAI Sora agreement (Disney)T1verified 2026.07.22“The Dec 2025 Disney-OpenAI deal, later collapsed in March 2026.”
- Sphere Wizard of Oz (Google)T1verified 2026.07.22“The 1939 Wizard of Oz expanded to the Sphere with Google generative models.”
New
The auteurs and the guilds weigh in
structuralA new Industry act opens with the directors' split (Scorsese joining Black Forest Labs, Cameron, Lucas, Jackson vs del Toro, Villeneuve, Spielberg, Nolan as DGA president) and the institutional rules (WGA/SAG-AFTRA/DGA deals, the Academy's AI eligibility rules, the White House letter).
Provenance
- Scorsese joins Black Forest Labs (Variety)T3verified 2026.07.22“Scorsese: 'Cinema is a young medium... we have to be open to how it can evolve'; used FLUX to storyboard his next film.”
- Nolan elected DGA President (DGA)T1verified 2026.07.22“Christopher Nolan elected DGA President, Sept 20, 2025, having chaired the AI Committee.”
- Cameron on AI and VFX cost (Variety)T3verified 2026.07.22“Cameron: cut VFX cost in half, 'not about laying off half the staff' but 'doubling their speed.'”
Version 15
July 22, 2026
New
AI filmmaking: vibe directing and agentic production
structuralA new part covers the turn from generating clips to directing films: vibe directing (OpenArt's June 2026 coinage, the film analogue of vibe coding), the agentic platforms led by Runway and OpenArt, the broader movement, and the authorship debate that dominated Cannes 2026.
Provenance
- OpenArt: what is vibe directingT1verified 2026.07.22“OpenArt coined vibe directing in June 2026: directing a film into existence by talking, shaping each shot conversationally.”
- Runway Gen-4.5 + General World ModelT1verified 2026.07.22“Runway's Gen-4.5 and first General World Model; the company reframes video as the prequel to simulating the world.”
- OpenArt Director (Hollywood Reporter)T3verified 2026.07.22“OpenArt's Director product and the vibe-directing framing.”
Upd
The book gets an act structure and a new title
structuralThe contents are regrouped into named acts (Foundations, The Big Picture, The Engines, The Craft, The Business, The Frontier, Reference), the glossary moves to the back, and the title broadens to reflect that the story is now the whole stack and AI filmmaking, not only the models.
Provenance
- Machine Cinema market mapT2verified 2026.07.22“The ecosystem view that motivated regrouping the book around the whole stack.”
Version 14
July 22, 2026
New
Realtime and interactive generation
structuralA different mode from batch: realtime canvases (Krea, Decart) and world models generate as fast as you move, collapsing the pipeline and folding finishing and distribution into the moment of generation.
Provenance
- Decart (Mirage realtime video)T1verified 2026.07.22“Sub-40ms/frame streaming diffusion for interactive, realtime video.”
- KreaT1verified 2026.07.22“Realtime generation canvas that renders as fast as the user works.”
New
Distribution and the microdrama economy
structuralDistribution is where the money lands. The vertical microdrama market (~$3.8B in-app in 2025, forecast to more than double in 2026) rivals Netflix for US mobile time; today it is mostly live action with AI as accelerant, but it is the format most exposed to full generation, with AI-native apps like Popshort and studios like Fairground emerging.
Provenance
- Deloitte TMT Predictions 2026 (short-form series)T2verified 2026.07.22“Micro-series in-app revenue ~$3.8B in 2025, forecast to more than double to ~$7.8B in 2026.”
- Sensor Tower: State of Short Drama Apps 2025T2verified 2026.07.22“ReelShort and DramaBox lead; Q1 2025 IAP ~$700M, ~3x YoY.”
- Fairground AI studio (Variety)T3verified 2026.07.22“Fairground: exclusively AI-generated series; founded by a Xumo co-founder.”
- YouTube generative tools + policy (DeepMind)T1verified 2026.07.22“Veo generation built into Shorts; synthetic content must be labeled and purely AI video is monetization-restricted.”
New
The audio pillar
structuralAudio (voice, dubbing, lip-sync, music, SFX) is nearly half the production phase and had been underweighted; a video pipeline without an audio plan is half a pipeline.
Provenance
- ElevenLabsT1verified 2026.07.22“Leading AI voice, dubbing, music, and sound-effects generation.”
- Machine Cinema market mapT2verified 2026.07.22“Voice, dubbing, lip-sync, music, and SFX are distinct first-class categories.”
New
The production pipeline as a chain
structuralA finished AI-native piece is a chain of specialized jobs (write, board, generate, hold consistency, voice, lip-sync, edit, grade, mark, distribute); generation is one link and not the hardest.
Provenance
- Machine Cinema market mapT2verified 2026.07.22“Each phase of the stack is a category because each is a distinct production job.”
New
The workflow and aggregator layer
structuralCreators increasingly work in multi-model canvases (Krea, Higgsfield, Flora, ComfyUI) rather than going model-direct; the workflow, not the model, owns the user, and agents are starting to drive the workflow.
Provenance
- Krea, unified multi-model canvas (Contrary Research)T2verified 2026.07.22“Krea: 20M+ users on a unified image+video canvas across many models.”
- Freepik/Magnific consolidation (Fortune)T3verified 2026.07.22“Magnific/Freepik: ~$200-230M ARR, bootstrapped and profitable, stack consolidation.”
New
The Stack: the book gets a new spine
structuralA new anchoring part reframes the field from the point of view of the whole creative stack rather than individual models: as generation commoditizes, value migrates up to workflows and down to distribution. Concrete backing: the application layer now captures the majority of enterprise gen-AI spend, and app-layer players like Higgsfield reached a ~$500M run rate in about a year.
Provenance
- Higgsfield $200M run-rate + Series AT1verified 2026.07.22“Higgsfield: 15M+ users, 4.5M generations/day, ~$200M ARR reported Jan 2026; ~70% enterprise.”
- Machine Cinema market mapT2verified 2026.07.22“The creative stack: development, production, post, inference, distribution.”
Version 13
July 22, 2026
New
When generation learned to reason
structuralA new Foundations chapter on reasoning-augmented image generation: GPT-Image-2's Thinking mode plans, web-searches, and self-checks before rendering (architecture undisclosed), alongside a genuine 2026 research revival of autoregressive image models.
Provenance
- Artificial Analysis (image)T2verified 2026.07.22“GPT Image 2 debuted #1 on the text-to-image arena, Elo ~1,338, ~+242 over #2 (largest gap on that board).”
- OmniGen-AR (arXiv)T2verified 2026.07.22
New
When video learned to talk
structuralA new Foundations chapter on native audio-visual generation: Veo, Kling 3.0's Omni Native Audio, and open-weight LTX-2 generate synchronized sound in the same pass as the picture; open models like Foley-Omni handle video-to-audio.
Provenance
- Google DeepMind VeoT1verified 2026.07.22
- Lightricks LTX-2 (open weights)T1verified 2026.07.22“Lightricks open-sourced LTX-2 on Jan 6, 2026: native synchronized audio + lip-sync, up to 4K/50fps, truly open weights.”
- Foley-Omni (arXiv)T2verified 2026.07.22
New
Provenance, watermarking, and the August 2026 deadline
structuralA new Operator's Playbook section on the EU AI Act Article 50 transparency obligations that apply from August 2, 2026 (deepfake disclosure and machine-readable marking), and the C2PA plus SynthID standards converging to meet them.
Provenance
- EU Commission: Article 50 transparencyT1verified 2026.07.22“Article 50 transparency obligations (deepfake disclosure + machine-readable marking) apply from 2 August 2026.”
- EU AI Act Article 50T1verified 2026.07.22
New
World models leave the lab
structuralA new section promoting world models from research to product: DeepMind's Genie 3 (real-time interactive worlds, ~1-minute memory, research prototype) and World Labs' Marble (a commercial paid product that exports Gaussian splats and meshes).
Provenance
- DeepMind GenieT1verified 2026.07.22“Genie 3: 720p, 20-24 fps, ~1 minute of memory; a limited-access research prototype, not exportable 3D.”
- TechCrunch (World Labs Marble)T3verified 2026.07.22
Upd
Agents drive the ComfyUI graph
structuralChapter 7 gains a subsection on Comfy MCP (launched June 30, 2026), the official bridge that lets AI agents assemble and run ComfyUI workflows on Comfy Cloud GPUs, plus in-canvas copilots.
Provenance
- Comfy MCPT1verified 2026.07.22“Comfy Org launched Comfy MCP on June 30, 2026; agents run workflows on Comfy Cloud GPUs.”
Upd
Blackwell economics and current pricing
detailThe economics chapter gains a hardware-and-pricing subsection: B200/GB200 rental ranges, Blackwell's inference-cost cut over Hopper, and current API reference prices (FLUX.2 from official docs; per-second video prices marked approximate).
Provenance
- B200 pricing (getdeploying)T3verified 2026.07.22
- FLUX pricing (official)T1verified 2026.07.22“BFL official docs: FLUX.2 [pro] from $0.03/image; [klein] ~$0.014-0.015.”
Version 12
July 22, 2026
Upd
ComfyUI deep dive: samplers, steps, and CFG scale
structuralChapter 7 gains a full treatment of the KSampler: what a denoising step does and where the returns flatten, how CFG scale trades prompt adherence against image quality, and how to pick a sampler and scheduler. The through-line is that settings must match the model class.
Provenance
- ComfyUI changelogT1verified 2026.07.22
Upd
OpenAI exits consumer video; Sora API sunsets September 24
detailThe Sora app went dark in April and the API will be discontinued September 24, 2026, amid roughly a million dollars a day in cost, a dead Disney deal, and copyright problems. Sora continues only as internal world-model research.
Provenance
- OpenAI HelpT1verified 2026.07.22“OpenAI: Sora app discontinued Apr 26, 2026; Sora API to be discontinued Sep 24, 2026.”
- The DecoderT3verified 2026.07.22
Upd
Google rebrands video as Gemini Omni
detailRather than a Veo 4, Google I/O introduced Gemini Omni, an any-to-any multimodal family. Gemini Omni Flash shipped the same day and sat at #1 across the major video arenas by July.
Provenance
- Google DeepMind VeoT1verified 2026.07.22
- Artificial Analysis (video)T2verified 2026.07.22
New
The Chinese video surge: Kling 3.0, Seedance 2.5, HappyHorse
structuralKling 3.0 Turbo added multi-shot prompting and native audio as Kuaishou raised at a reported ~$18B valuation; ByteDance's Seedance 2.5 generates native single-pass 30-second 4K clips; Alibaba's HappyHorse topped the open video arenas with single-pass joint audio and video.
Provenance
- Kuaishou IRT1verified 2026.07.22
- TechTimes (Seedance 2.5)T3verified 2026.07.22
- fal (HappyHorse)T2verified 2026.07.22
New
Meta enters image generation with Muse
structuralMeta Superintelligence Labs launched Muse Image, its first image model, an agentic system that plans a layout before rendering, shipping free inside the Meta AI app, Instagram, and WhatsApp.
Provenance
- Meta AI blogT1verified 2026.07.22
- Meta NewsroomT1verified 2026.07.22
New
The open-weights wave: Ideogram 4.0, Krea 2, HiDream-O1
structuralThree serious open-weight image models shipped in five weeks, closing the gap on the closed flagships: Ideogram 4.0 with structured-JSON prompting, Krea 2 with a two-second Turbo variant, and the MIT-licensed pixel-native HiDream-O1.
Provenance
- Ideogram 4.0T1verified 2026.07.22
- Krea newsT1verified 2026.07.22
- HiDream-O1T1verified 2026.07.22
Upd
Google retires Imagen for the Gemini image line
detailGoogle began sunsetting the Imagen family, with API endpoints shutting down through August 2026 and the migration path pointing to Nano Banana. Nano Banana 2 Lite became the fastest, cheapest tier at ~4s and ~3 cents per image.
Provenance
- Gemini API deprecationsT1verified 2026.07.22
- TechCrunch (Nano Banana 2 Lite)T3verified 2026.07.22
Upd
The tier below Google and OpenAI is a race
detailRecraft V4.1 briefly led outside Google and OpenAI in May, but by late July had been passed by Reve 2.1 and Microsoft's MAI-Image-2.5 on the Artificial Analysis arena.
Provenance
- Artificial Analysis (image)T2verified 2026.07.22
New
New frontier entrants: Reve 2.1, MAI-Image-2.5, Grok Imagine 1.5
detailReve 2.1 and Microsoft's MAI-Image-2.5 now rank #2 and #3 on the Artificial Analysis image arena, and xAI's Grok Imagine Video 1.5 (May 31) briefly topped the image-to-video arena with native audio. All three were absent from earlier editions.
Provenance
- Artificial Analysis (image)T2verified 2026.07.22“As of late July, Reve 2.1 (#2) and Microsoft MAI-Image-2.5 (#3) rank above Recraft.”
- xAI Grok Imagine 1.5T1verified 2026.07.22