The Stack: Products, Workflows, and the Whole Board
The audio pillar
Chapter 92
2 min read
Reviewed v78 · August 2026
The audio pillar
For most of its length this book has stayed with the visual, image and video, a little too literally. Look at the map and audio is nearly half of the production phase. A silent guide to a medium that now speaks is incomplete.
Audio in 2026 splits into a few real categories. Voice and dubbing, led by ElevenLabs and a deep field including Play HT, Cartesia, Hume, and Respeecher, generate and clone speech and localize it across languages. Lip-sync and avatars, from HeyGen and Synthesia to Hedra and Sync, marry a voice to a face. Music generation, led by Suno and Udio, produces full songs from a promptPromptThe text description you provide to a model to specify what you want it to generate.. Sound effects and foleyFoleyThe everyday sound effects of a scene, footsteps, cloth, impacts, now increasingly generated from the video itself., from ElevenLabs and open models such as MMAudio and Foley-Omni, generate the ambient and impact sound that makes a scene feel real.
Two shifts matter. The first, covered in the foundations, is that frontier video models now generate synchronized audio natively, so the line between the video pillar and the audio pillar is blurring at the top. The second is that audio has its own reliability and rights problems, voice cloningVoice cloningReproducing a specific person's voice from a short sample so it can speak new lines, often in any language. most of all, which is exactly why the likeness-protection tools in the post-production phase exist.
For an operator the practical point is that a video pipeline without an audio plan is half a pipeline. Dialogue, music, and sound are not garnish, they are most of what makes generated footage read as finished rather than as a demo. Budget for them, tool for them, and decide early whether your video model gives you audio natively or whether you are stitching it in.
The reason audio earns its own pillar rather than a footnote is that it is where the illusion of a finished piece is won or lost. A visually flawless clipCLIPA text encoder developed by OpenAI in 2021 that learns to align text and images in a shared embedding space. Foundation of most text-to-image models from 2022 onward. with wrong or missing sound reads as a tech demo; a rougher clip with convincing voice, room tone, and a score reads as a film. The stack here has three layers that mostly still run as separate tools. Speech, where ElevenLabs and a deep field of cloning and dubbing models put words in a character's mouth in any language. Music, where Suno and Udio generate a scored track from a line of description. And sound effects and foley, where the newer video-to-audioVideo-to-audioGenerating a soundtrack timed to an existing video clip, so effects land on the exact frame of the on-screen action. Also called video-to-sound. models watch the picture and place the footsteps and impacts on the action rather than near it. The shift to watch is that the frontier video models have started generating synchronized audio in the same pass, which will eventually pull some of this pillar back into the video box, but for now assembling the soundtrack from specialists is still the norm and still where a lot of the craft quietly hides.
Check your understanding
pass: 5 of 7
Answer at least 5 of 7 correctly to unlock the next chapter.
1. Which audio category are Suno and Udio associated with?
2. What does the lip-sync and avatars category do?
3. How is the line between the video and audio pillars blurring at the top?
4. Why do likeness-protection tools exist in the post-production phase?
5. What is the main practical point for an operator about audio?
6. What decision should an operator make early regarding audio?
7. Which tools are named for the voice and dubbing category?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.