Contents

108 / 153

The Craft of Generation

Directing audio: voice, music, and sound

Chapter 107

2 min read

Reviewed v78 · August 2026

Getting a believable performance out of an audio model is a craft, and it transfers from the craft you already know. Prompting a voice model is closer to directing an actor than to typing a command, prompting a music model is closer to briefing a composer, and prompting for sound effects is closer to spotting a film. In each case you supply intent the model cannot infer, then iterate on takes.

For voice, direct three things: emotion, pacing, and pronunciation. Newer models accept explicit performance direction, so an instruction to deliver a line calmly, or with a hint of reluctance, changes the read. Control pacing with punctuation and paragraphing, since commas, periods, and line breaks are the levers for pauses and breath. Fix pronunciation deliberately rather than hoping, spelling tricky names phonetically and using the tool's pronunciation controls where offered. When you clone a voice, the input clip is the performance ceiling, so use a clean, dry reference with no music or reverb and a consistent tone.

For music, brief the model the way you would brief a session: genre, mood, instrumentation, tempo, and structure. Vague prompts yield generic beds, so be specific about the arc you want, an intro that builds, a chorus that lifts, a quiet bridge. Then iterate, keeping what works and re-rolling sections rather than starting over. The real editing power comes from stems: generating separate vocal, drum, bass, and melody tracks lets you rebalance, replace, or extend parts in a digital audio workstation instead of accepting the mix as fixed, which is how AI music actually enters a professional pipeline.

For sound effects and foley, the craft is timing to on-screen action. Text-to-sound turns a description into an effect with control over duration and style, while video-to-sound analyzes footage and proposes effects that match events, engine roars and tire screeches over a car chase. Whichever path you take, the work is synchronization: place hits on the frame, layer ambience under discrete effects, and check against picture, because a sound that is plausible but a beat late still reads as wrong. Video-native models like Veo 3 now generate synchronized dialogue, effects, and ambience in one pass, which shifts some of this work from generation toward selection and cleanup.

Check yourself0 / 4

Q01

What three things should you direct when prompting a voice model?

Q02

What is the two-consents rule for voice cloning?

Q03

Why do stems give the real editing power in AI music?

Q04

What is the core craft when adding sound effects and foley?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.