Contents

112 / 153

AI Filmmaking: Vibe Directing and Agentic Production

The craft: consistency and control

Chapter 111

3 min read

Reviewed v78 · August 2026

Character consistency is the problem the field spent years failing to solve, and the current everyday answer is reference-image conditioning. Runway's Gen-4 (March 2025) lets you supply reference images of a character or scene and generate consistent versions across new shots, angles, and lighting with no per-character training, and Google's Veo 3.1 (October 2025) added the same idea under the name reference images, or ingredients, accepting a few references of a character, object, or scene to steer generation. The workflow is now: lock identity in one strong reference, then reuse it for every shot. When a reference is not enough for a lead who must survive dozens of scenes, pros still train an identity LoRA, a small model fine-tuned on one subject. One published method (from the AI-film school Curious Refuge) is concrete: collect at least ten varied images of the subject, avoid multiple frames from the same shoot, keep faces unobscured, use the character's name as the trigger word, and train roughly a thousand steps, a job of minutes on a hosted service. Treat those numbers as one credible recipe, not a universal standard.

Motion and camera are directed two ways, through the driving image and through language. Camera-control prompting, the directorial vocabulary of dollies, cranes, whip pans, push-ins, dolly zooms, and rack focus, is now understood well enough that tools market those moves as callable from the prompt. This is where a real cinematography vocabulary pays off literally: the knowledge that lets a director say slow push-in on a fifty millimeter, shallow depth of field is the same knowledge that steers the model, which is why directing language transfers almost verbatim. First-and-last-frame keyframing is the strongest form of control: you provide the opening and closing image and the model synthesizes the transition, turning an unpredictable generation into a bounded interpolation that starts and ends where the edit needs it. Shot extension addresses the length ceiling by generating clips that connect to a previous one, so continuity is engineered clip by clip rather than captured in one take.

Performance is now its own craft stage. Runway's Act-One (October 2024) drives a generated character from a single actor's video and voice, preserving eye-lines, micro-expressions, pacing, and delivery, from a simple single-camera setup with no motion-capture rig or manual face-rigging, and Act-Two (July 2025) extended capture to full head, body, and hand gestures from a driving video plus a reference character. This is modern puppeteering: the actor's job survives, the mocap stage and the rig do not. It is also why the strongest AI films route their most demanding moments through a real human performance rather than trying to synthesize emotion from text.

Check yourself0 / 6

Q01

Why do pros work image-first, locking a still before animating?

Q02

What is the everyday answer to character consistency, and when do pros still go further?

Q03

The chapter says identity drift is reduced, not solved. Where does it still break, and what drifts fastest?

Q04

Why does a real cinematography vocabulary pay off literally in AI filmmaking?

Q05

Why is first-and-last-frame keyframing described as the strongest form of control?

Q06

What is the 'hard wall' the models still hit, and what is the practical response?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.