Contents

113 / 153

AI Filmmaking: Vibe Directing and Agentic Production

The craft: technique, volume, and the edit

Chapter 112

3 min read

Reviewed v78 · August 2026

If there is a line between AI work that looks finished and AI work that looks like slop, most of it comes down to a handful of techniques the experienced use and beginners skip, plus a set of working habits that turn out to be remarkably consistent across everyone doing it well.

Consistency is the first battleground, and the two dominant methods are reference sheets and trained identities. A character sheet is a set of multi-angle reference images of a face and costume, fed into each shot so the model has something to hold to. A character LoRA goes further, training a small adapter on a set of images so a fixed identity can be summoned across an entire film. Runway added a casting feature that keeps an AI actor visually consistent across a production, and the image and video labs increasingly expose reference-based controls for the same purpose. Where a model exposes its random seed, reusing that seed across related shots is a cheap way to hold appearance steady.

The second battleground is control. Camera-control prompting, the directorial vocabulary of dollies, cranes, whip pans, and pushes to a close-up, is what separates a shot that feels shot from a shot that merely moves. Keyframing the first and last frame pins a shot's composition, shot extension chains short clips into continuous motion, and a dedicated upscaling and grading pass at the end is what takes a generation from demo to deliverable. None of this closes the gap entirely. Sustained consistency over a long runtime, dialogue-heavy scenes with several characters, fine physical continuity, and long controllable takes remain genuinely hard. The tools have not made the hard parts easy. They have made them possible for a small team, which is a different and more interesting thing.

Underneath the techniques is a common stack and a set of habits, and they are the real curriculum. The stack is strikingly uniform: a language model to write the script, shot list, and prompts; a still-image model for art-directable keyframes; an image-to-video model to bring those frames to motion; an upscaler to finish; a normal editor to assemble; and separate tools for voice and music. Image first, because a still can be art-directed in a way raw text-to-video cannot.

Then the habits. The first and most important is volume followed by a ruthless cull. The ratios are startlingly consistent: something near a few hundred generations for every used shot is normal. PJ Accetturo generated three to four hundred clips for the roughly fifteen that made his Kalshi ad. Paul Trillo kept fifty-five shots out of seven hundred for his Sora music video. The dabbler generates once and is disappointed. The professional generates a hundred times and curates. The second habit is no tool loyalty: pros mix models for their strengths, one engine for realism and native audio, another for stylized motion, another for character reference, a fourth to upscale, with allegiance to none. The third, and the one every experienced creator stresses, is that the craft lives in the edit. The magic is not a single stunning generation, it is the person who knows where to cut and when to let a real voice and human music do the emotional lift.

The discipline is measurable at the extreme. In one documented animated-episode case, a creator generated 164 clips for a three-minute cut, used 41, kept an average of about five seconds from each fifteen-second clip, and built seventeen of the forty-one final shots by stitching more than one generation into a single composite. The traditional shooting ratio is not just inverted, it is inverted many times over.

Fig.diagram
GENERATEDCULLby tasteKEPTa handfulPJ Accetturo: about 350 clips to 15. Paul Trillo: about 700 to 55.The shooting ratio, inverted.Overgeneration is a planned budget line, not waste.The scarce skill is knowing which few seconds of a clip actually work. The craft lives in the cull.
The shooting ratio, inverted, hundreds of generations culled to a handful of keeps.

Finishing still happens in traditional software, and this is where amateur work gives itself away. Upscaling and restoration tools like Topaz clean, denoise, and enlarge footage before the edit, and a normal grade in DaVinci Resolve is what makes disparate generations feel like one film. The practical move is to pick a hero shot, grade it, then match every other clip to it on the scopes, so generations produced under inconsistent model lighting resolve into a single intentional look. The generative tools produce raw material. The finish is still human, and still traditional.

Check yourself0 / 6

Q01

The chapter says the tools did not make the hard parts easy. What did they actually do?

Q02

What is the single most important working habit, and roughly what ratio illustrates it?

Q03

What does 'no tool loyalty' mean in practice for AI filmmakers?

Q04

Why is the traditional shooting ratio described as inverted, and what mindset fights the medium?

Q05

Why does the chapter say the craft ultimately lives in the edit?

Q06

What is the practical grading move that makes disparate generations feel like one film?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.