Almost all of the AI that reaches a real screen does not make the whole film. It enters a live-action production at a few specific, expensive, well-bounded handoffs and leaves the rest of the pipeline alone. This is the hybrid model, and it, not the fully generated feature, is where professional adoption actually lives. The shoot stays human: real actors, a real camera, a director and a cinematographer making a take. AI arrives as an accelerant at the edges. It generates or automates a draft layer, then a supervisor scopes it, corrects it, and owns the final frame at delivery resolution.
The cost pressure behind all of this is real. Los Angeles logged 19,694 on-location shoot days in 2025, down about sixteen percent from the year before, and scripted television, the industry's former backbone, has fallen from a 2021 peak of 18,560 shoot days to 7,716 in 2024, a drop of nearly sixty percent in three years. Production chased tax credits to Georgia, the United Kingdom, and Eastern Europe, and California expanded its own film incentive in mid 2025 to fight back. Hybrid production is one answer to that squeeze, because it can put the scale of a big film onto a local stage at a fraction of the location and build cost.
And hybrid has two halves, not one. The back half is the post-production work, where generated shots are finished by hand. The front half happens on the stage, and it predates the current models: virtual production, where actors perform in front of a large LED volume showing a real-time rendered world that also lights them, the method The Mandalorian made mainstream with Industrial Light and Magic's StageCraft, and performance capture, where a real actor drives a digital character, the lineage Avatar advanced. Generative AI now accelerates both by creating the environments the volume displays and the assets a crew would otherwise build by hand. The team behind Prime Video's House of David took exactly this route and went on to launch Innovative Dreams with Luma, a venture built around what it calls real-time hybrid filmmaking, blending performance capture, virtual production, and AI on one stage.
This is already how the industry works. On its July 2026 earnings call Netflix said roughly three hundred titles used generative AI across the year, spanning previsualization, visual effects, relighting, and crowd simulation, and it cited one project whose AI-enhanced footage was produced about twice as fast at about half the cost. Its 2025 series El Eternauta was its first generative-AI final footage on screen, a building collapse it said finished roughly ten times faster and became feasible on a budget that could not have afforded conventional visual effects. The work is concentrated in post, exactly where the hybrid thesis says it should land. What follows is a working guide to running that process, from the first planning decision to the final insured delivery. The through line never changes: AI drafts a layer, a human scopes and finishes it, and the film stays something you can own, clear, and insure.
Decide what is filmed and what is generated
The first decision, made in prep, is a division of labor across the frame. Get it wrong and you either shoot something a model could have made for a fraction of the cost, or you ask a model to carry a shot it cannot hold. A workable rule: film anything the audience reads as a real human performance, film faces in sustained close-up, and film anything that must be legally cleared as a real person's likeness. Lean on AI for what is expensive, repetitive, or physically impossible: crowd tiles, set and horizon extensions, distant doubles, de-aging, weather and time-of-day changes, and the single impossible shot a budget could not otherwise afford.
A sharper version of the same test asks where the shot sits in the pipeline. AI is strongest at the front (concepting and previs) and at the back (cleanup, de-age, extensions, finishing), and weakest in the middle, the sustained, character-driven performance take where identity, continuity, and subtle acting have to hold for a full beat. Use AI to prepare a shot and to repair a shot, and keep the human performance in between.
There is a mechanical reason the middle resists generation. On human subjects, pure generative output has three persistent tells. It lacks physical weight, since real people shift their balance and carry the mass of what they touch while generated figures tend to drift and float. Its texture is too perfect, missing the moisture, grit, and optical flaws a real lens records. And it degrades over time, because the longer a synthetic shot holds, the more small anomalies the eye accumulates. All three get worse exactly where a performance lives, in the sustained close-up, which is the shot you photograph.
Previs: plan the shot, do not commit it
In development, use models to see the film before you can afford to shoot it. Generate concept frames, extend a keyframe into a rough moving sequence, and cut a quick animatic so the director, the cinematographer, and the department heads argue over real images instead of words. Studios are wiring this in directly: Runway signed a model-training deal with Lionsgate in September 2024 aimed at storyboarding, editing, and post. The payoff is speed, a look that took weeks of storyboarding can land in days.
The discipline is to treat previs as a conversation, not a plan of record. Generated previs is not shot accurate: the lens, the set dimensions, the eyelines, and the continuity are approximations, so a previs supervisor still has to translate an approved look into a real camera and lighting plan. Lock the story, the staging, and the intent from the previs, then re-derive every technical number on the day. The animatic tells you what the shot feels like, not what focal length delivers it.
On the stage: the LED volume and performance capture
Before a single frame is generated, the other half of hybrid happens physically, on a stage. Virtual production replaces the green screen with a large LED volume, a curved wall and ceiling of LED panels that display a real-time rendered environment running in a game engine. The Mandalorian built the mainstream template with Industrial Light and Magic's StageCraft: actors performed inside a roughly twenty-foot-high, two-hundred-seventy-degree wall on a seventy-five-foot stage, with the environment rendered live in Unreal Engine, and more than half of the first season was shot that way. The win is not only the background, it is the light. The wall throws real, correct illumination and reflection onto faces, skin, and chrome, which a green screen never can.
Performance capture is the other practical pillar. A real actor's face and body are recorded and mapped onto a digital character, so the performance, the timing, the micro-expressions, and the weight stay human even when the body on screen is not, the lineage Avatar pushed into the mainstream. Generative AI slots into both methods rather than replacing them: it can generate the environment the LED volume displays, extend or restyle it on the fly, and produce the background assets and creatures a team would otherwise model by hand. The stage keeps the performance and the physics real, and AI supplies the scale.
On set: shoot for the hybrid finish
Everything AI does in post is only as good as the plate it is handed, so a hybrid shoot is a data-capture job as much as a performance one. You shoot the scene, then you protect the AI work with a few disciplined extra passes. The cost of skipping them is not visible on the day, it shows up as a doubled post bill weeks later.
Clean plates. Grab a locked-off or emptied pass of the shot so post has a clean background to paint into, extend, or remove an element from. A paint-out with a clean plate is minutes; without one it is hours of invented detail.
A lighting and lens reference. Shoot a grey ball, a chrome ball, and a quick reference frame, and note the lens and camera data, so any generated or relit element can be matched to the real light instead of guessed. AI relighting tools work far better when they have the real lighting to match.
Tracking markers and coverage. Put tracking markers on anything that needs a camera solve or a CG double, and shoot the actor's real eyeline and interaction even when the other element will be added later. A model can extend a set, but it cannot invent a performance the actor never gave to empty air.
A real-time reference where it earns its place. A live composited or de-aged monitor feed helps the director and the actors judge framing and performance. On Robert Zemeckis's Here, Metaphysic Live rendered de-aged faces on the leads' live performances and fed the director a monitor at roughly a six-frame delay plus a youth mirror for the actors. Treat that image the way it was treated there: a reference to guide the take, never the pixel that ships.
Consent and likeness capture, on the day. If a shot involves a digital replica, a de-age, a double, or a voice, capture the scan, the reference, and the signed use terms while the performer is present and paid. It is far cheaper and cleaner than reconstructing consent after the edit locks.
Then hand post a package, not just footage: the plates, the reference passes, the camera and lens data, and the consent paperwork. A missing clean plate or a lost lens number is what turns a two-hour fix into a two-day one.
In post: the AI-assisted effects loop
Post is where most of the AI actually lives, and it runs as a loop rather than a single pass. For each shot the sequence is the same: pick the task, generate a draft layer, integrate it, and finish it by hand.
Map the task to the tool. For faces, de-aging, and digital doubles, Metaphysic (now DNEG's Brahma division) and MARZ's Vanity AI for 2D de-age and beauty at shot scale. For in-plate edits, object and rig removal, paint-out, set changes, relighting, and new camera angles, Runway's Aleph. For dropping a tracked CG character over a live actor with automatic clean plates and mattes, Autodesk's Flow Studio. For relighting flat footage after the fact, Beeble's SwitchLight. The right tool is the one that matches the specific task, not one model asked to do everything.
DIY or vendor is a scale decision. A team of one can run these tools directly for a short film or a commercial and get a broadcast-clean result. A feature routes the same tools through a VFX house, because the hard part is not generating one effect, it is holding it across hundreds of shots at delivery quality with someone accountable for the result.
Techniques that sell the composite
Getting a generated element to live inside a real plate is its own craft, with a few repeatable moves that do most of the work. They are the difference between a shot that reads as cinema and one that reads as a demo.
Keep the synthetic shot short. Suspicion scales with screen time, so the longer a fully generated shot holds, the more time the eye has to find the tells. Sandwich it instead: cut from a real, practically lit shot into a brief generated one, a second or two, then back to a real anchor. The authentic frames set a baseline of physical reality that the short insert borrows, and the viewer's radar never trips. It is also why fast-cut forms like microdrama and music video hide AI better than a long unbroken take.
Anchor the wide shot in something real. The most convincing impossible environments are grown from a real plate, not generated whole. Shoot the actor doing the real thing in a small controlled setup, a few feet of water on a stage, a corner of a set, a patch of ground, then feed that frame to a generative tool and extend the borders outward into the vast world while holding the real center as the anchor. The eye locks onto the real geometry, lighting, and performance in the middle and accepts the generated surroundings. A confined stage becomes an ocean, and the part that had to feel alive was photographed.
Bridge two real frames with a generated move. When a camera move is physically impossible or simply unaffordable, shoot the two real endpoints, a start frame and an end frame matched for light and composition, and let a video model synthesize the motion between them. The book covers this as first-and-last-frame keyframing, and it is among the strongest hybrid moves precisely because both anchors are real and only the transition is invented.
Finishing and localization
Two more stages sit at the very end. In finishing, tools like Topaz Video upscale, denoise, and interpolate archival or mismatched footage toward delivery resolution, with the colorist watching for fabricated detail that was never in the source. Use it to rescue a noisy plate or match resolutions across cameras, not to invent texture that changes what the shot is.
In localization, AI turns one shoot into many market versions. Audio dubbing re-voices dialogue across languages (ElevenLabs's Dubbing v2 claims a hundred and seventy five), while visual dubbing re-animates the on-screen actor's mouth to fit the new language (Flawless's TrueSync). The same family of tools refines dialogue in the original language too: The Brutalist used Respeecher to perfect the leads' Hungarian pronunciation of their own recorded lines, and the disclosure alone set off an awards-season debate even though the performances were entirely the actors' own. Premium work still routes all of it through human translators, voice directors, mixers, and a QC pass, and the disclosure question is now part of the job.
Clear it: consent, rights, and insurance
The step that separates a shippable hybrid film from an experiment is clearance, and it starts in prep, not in the edit. Three things have to be true at delivery, and each one is easier to secure on the day than to reconstruct later.
Consent for every likeness. Any digital replica, a face swap, a voice clone, a de-age, or a full digital double, needs the performer's informed, written consent with a specific description of the use, and reuse on a different project needs fresh consent and additional pay. This is the core of SAG-AFTRA's digital-replica framework. Capture it while the performer is present and keep it filed with the asset it covers.
A clean chain of title. Every element in the frame must be created, licensed, or cleared, and you must be able to prove where each one came from. This is the concrete reason a fully generated shot built on a model trained on unlicensed data is dangerous: you cannot show its provenance. Keep the source and rights for generated elements the way you keep releases for locations, music, and extras.
Insurable and copyrightable at the end. Distributors require errors and omissions coverage to carry a film, and that underwriting demands the clean chain of title above. The US Copyright Office requires human authorship, so the live-action performance, the cinematography, and the direction are what make the finished film registrable and defensible. Hybrid clears all three gates. Fully generated clears none of them cleanly. That is the whole financial case for keeping a human production around the AI.
Common mistakes, and a working checklist
The failures repeat, and they are all avoidable in prep. Generating what you should have shot, asking a model to carry a real performance it cannot sustain. Shooting without the plates, no clean pass, no lighting reference, no tracking, so post pays triple to reverse-engineer them. Skipping consent on the day and trying to paper it over after the edit locks. Budgeting the generation but not the finish, then discovering the eighty-percent draft still needs a compositor. Treating previs as a plan and shooting to lens and staging numbers that were never real, and letting a synthetic take run long when a one or two second insert between real anchors would have passed unnoticed.
A working sequence, prep to delivery: divide the frame into filmed and generated in prep; previs to align the team, then re-derive the real camera plan; on the day shoot the performance plus clean plates, a lighting and lens reference, tracking, and consent; in post pick the tool per task, generate a draft, and finish it to 4K with a supervisor accountable; then upscale, color, and localize; and clear consent, chain of title, and insurance before you deliver. Run it in that order and AI makes the film cheaper and more ambitious without ever making it something you cannot own.