Getting a believable performance out of an audio model is a craft, and it transfers from the craft you already know. Prompting a voice model is closer to directing an actor than to typing a command, prompting a music model is closer to briefing a composer, and prompting for sound effects is closer to spotting a film. In each case you supply intent the model cannot infer, then iterate on takes.
For voice, direct three things: emotion, pacing, and pronunciation. Newer models accept explicit performance direction, so an instruction to deliver a line calmly, or with a hint of reluctance, changes the read. Control pacing with punctuation and paragraphing, since commas, periods, and line breaks are the levers for pauses and breath. Fix pronunciation deliberately rather than hoping, spelling tricky names phonetically and using the tool's pronunciation controls where offered. When you clone a voice, the input clipCLIPA text encoder developed by OpenAI in 2021 that learns to align text and images in a shared embedding space. Foundation of most text-to-image models from 2022 onward. is the performance ceiling, so use a clean, dry reference with no music or reverb and a consistent tone.
For music, brief the model the way you would brief a session: genre, mood, instrumentation, tempo, and structure. Vague prompts yield generic beds, so be specific about the arc you want, an intro that builds, a chorus that lifts, a quiet bridge. Then iterate, keeping what works and re-rolling sections rather than starting over. The real editing power comes from stemsStemsThe separated individual tracks of a piece of music, such as vocals, drums, and bass, which can be edited or remixed on their own.: generating separate vocal, drum, bass, and melody tracks lets you rebalance, replace, or extend parts in a digital audio workstation instead of accepting the mix as fixed, which is how AI music actually enters a professional pipeline.
For sound effects and foleyFoleyThe everyday sound effects of a scene, footsteps, cloth, impacts, now increasingly generated from the video itself., the craft is timing to on-screen action. Text-to-sound turns a description into an effect with control over duration and style, while video-to-sound analyzes footage and proposes effects that match events, engine roars and tire screeches over a car chase. Whichever path you take, the work is synchronization: place hits on the frame, layer ambience under discrete effects, and check against picture, because a sound that is plausible but a beat late still reads as wrong. Video-native models like Veo 3 now generate synchronized dialogue, effects, and ambience in one pass, which shifts some of this work from generation toward selection and cleanup.
Check yourself0 / 4
Q01
What three things should you direct when prompting a voice model?
Q02
What is the two-consents rule for voice cloning?
Q03
Why do stems give the real editing power in AI music?
Q04
What is the core craft when adding sound effects and foley?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.