Contents

93 / 153

The Stack: Products, Workflows, and the Whole Board

The audio pillar

Chapter 92

2 min read

Reviewed v78 · August 2026

For most of its length this book has stayed with the visual, image and video, a little too literally. Look at the map and audio is nearly half of the production phase. A silent guide to a medium that now speaks is incomplete.

A mixing console. Audio is nearly half of a finished piece, and still mostly its own craft.
A mixing console. Audio is nearly half of a finished piece, and still mostly its own craft.Mixing on a Rupert Neve Designs 5088 console by VACANT FEVER, CC BY-SA 2.0, via Wikimedia Commons

Audio in 2026 splits into a few real categories. Voice and dubbing, led by ElevenLabs and a deep field including Play HT, Cartesia, Hume, and Respeecher, generate and clone speech and localize it across languages. Lip-sync and avatars, from HeyGen and Synthesia to Hedra and Sync, marry a voice to a face. Music generation, led by Suno and Udio, produces full songs from a prompt. Sound effects and foley, from ElevenLabs and open models such as MMAudio and Foley-Omni, generate the ambient and impact sound that makes a scene feel real.

Two shifts matter. The first, covered in the foundations, is that frontier video models now generate synchronized audio natively, so the line between the video pillar and the audio pillar is blurring at the top. The second is that audio has its own reliability and rights problems, voice cloning most of all, which is exactly why the likeness-protection tools in the post-production phase exist.

For an operator the practical point is that a video pipeline without an audio plan is half a pipeline. Dialogue, music, and sound are not garnish, they are most of what makes generated footage read as finished rather than as a demo. Budget for them, tool for them, and decide early whether your video model gives you audio natively or whether you are stitching it in.

The reason audio earns its own pillar rather than a footnote is that it is where the illusion of a finished piece is won or lost. A visually flawless clip with wrong or missing sound reads as a tech demo; a rougher clip with convincing voice, room tone, and a score reads as a film. The stack here has three layers that mostly still run as separate tools. Speech, where ElevenLabs and a deep field of cloning and dubbing models put words in a character's mouth in any language. Music, where Suno and Udio generate a scored track from a line of description. And sound effects and foley, where the newer video-to-audio models watch the picture and place the footsteps and impacts on the action rather than near it. The shift to watch is that the frontier video models have started generating synchronized audio in the same pass, which will eventually pull some of this pillar back into the video box, but for now assembling the soundtrack from specialists is still the norm and still where a lot of the craft quietly hides.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. Which audio category are Suno and Udio associated with?

  2. 2. What does the lip-sync and avatars category do?

  3. 3. How is the line between the video and audio pillars blurring at the top?

  4. 4. Why do likeness-protection tools exist in the post-production phase?

  5. 5. What is the main practical point for an operator about audio?

  6. 6. What decision should an operator make early regarding audio?

  7. 7. Which tools are named for the voice and dubbing category?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.