Two companies show the two routes into agentic filmmaking, and a wider field is filling in around them.
Runway took the model-first route and kept climbing. It began as a video-generation lab and built outward into a filmmaking platform. Gen-4, released in March 2025, was pitched as the first Runway model to hold characters, objects, and environments consistent across shots rather than treating each frame independently. Aleph, in July 2025, added in-context video editing, so you bring a shot, describe the change, and it re-renders coherently, adjusting shadows and lighting to match. Act-Two, later that July, drives a synthesized character performance from a reference image plus a driving video. Alongside the tools, Runway began framing itself beyond film, presenting video generation as a step toward simulating the world. Runway is at once a filmmaker's tool and a bet that the same technology becomes a general world simulator.
OpenArt took the workflow-first route and arrived at the same place from the other side. Rather than train a flagship model it orchestrates several, Runway among them, behind a conversational director. Its Director product, launched on June 26 2026, generates films of up to about five continuous minutes with synchronized visuals, voice, music, and sound effects, where rival tools still think in five-to-fifteen-second clips. It is the product that turned vibe directing into a phrase. Where Runway sells you the best engine, OpenArt sells you the director's chair over everyone's engines.
It is worth going one level deeper into how the loop actually runs, because the mechanics are where the craft lives and where the honest limits show. OpenArt's Director walks you from idea to story to a scene and shot breakdown, then to images, clips, voiceover, and music, holding faces, voices, and style consistent and blending several underlying models. Its real differentiator is revision: change one shot with a note and it re-renders that shot without re-rolling the film, which is what makes directing it feel like giving notes to an editor. Its weak link is writing. First-pass scripts are rudimentary and need real rewriting, so the human's job concentrates exactly where you would expect, on story.
The others sort along the same axis. Google's Flow is built around Veo and organized for assembly rather than single clips, with camera controls that set angle and motion by description, a timeline for sequencing and extending shots, and a reference system for reusing characters and objects across shots. Lightricks' LTX Studio comes at it script first, auto-breaking a screenplay into scenes and shots and extracting reusable characters, aimed at producers doing previs and pitch decks rather than final frames. Fable's Showrunner is the outlier, offering viewer-steered animated episodes you can prompt into existence, genuinely novel but animation only and, by the honest account of people who have watched, still rough.
The pattern across all of them is the same. The tools are getting very good at assembly, consistency, and control, and they remain weakest at writing and at anything that has to hold up as live-action drama. That is not a knock, it is a map. It tells you where the software carries the load and where the human still has to.
Look closer at each platform and the split gets concrete. Runway is model-first and keeps bolting orchestration on top. Gen-4 (March 2025) is built around world consistency, generating a character or location across new angles and lighting from reference images with no per-character training. Gen-4 Aleph added video-to-video editing in context, where you supply footage and describe a change (relight, change the weather, add or remove an object, generate a new camera angle) and it edits inside the original scene. In December 2025 Runway pushed to native audio and dialogue with roughly one-minute takes, and shipped its first world model, a system that generates interactive scenes rather than fixed clips. The world-simulation framing is the strategic tell: Runway is betting that predicting pixels forward in time, not just rendering prettier frames, is the path to control.
OpenArt's Director attacks the opposite end, the revision problem, and coined the phrase vibe directing. It produces pieces up to about five continuous minutes, holds characters, faces, voices, and products consistent from first scene to last, and supports several languages with lip-sync. The claim that matters is the loop, not the length: you leave a persistent note (change the ending) and only that part re-renders. Its honest weakness, confirmed by early users and the reporters who covered the launch, is writing: first drafts are rough and need real rewriting, so the human's job concentrates on story. Google's Flow is the platform-scale entry, custom-built for Veo, with Camera Controls for motion and angle, Scenebuilder to extend shots into sequences, and Ingredients to reuse characters and objects across clips. Lightricks' LTX Studio is the script-first previs specialist that breaks a screenplay into scenes and shots for pitch decks rather than final frames. And Higgsfield and Krea are openly orchestration-first aggregators, routing across Kling, Veo, Sora, Seedance, Wan, and more, on the bet that model leadership rotates every few months so the durable value is the workflow around whichever model is best this quarter.