Contents

59 / 153

Audio Models: Voice, Music, and Sound

Native audio in video, and the pressure it creates

Chapter 58

3 min read

Reviewed v78 · August 2026

The single biggest structural force in audio right now is not a new audio model at all. It is that the video models have started generating their own sound. Until 2025, AI video was effectively silent film: you rendered the picture, then found or made audio separately and tried to line it up. Native audio changes the unit of production. The model emits picture and sound together in one pass, so footsteps land on the frame the foot touches, dialogue tracks the lips, and room tone matches the room. Because the two signals are generated jointly rather than stitched together, synchronization is built into the render rather than bolted on afterward.

Google's Veo 3 was the first major model to ship this, announced at Google I/O in May 2025, generating dialogue, sound effects, and ambience natively alongside the video. OpenAI's Sora 2 followed on September 30, 2025, producing synchronized dialogue, effects, and soundscapes inside the same generation pipeline. Through 2026 the rest of the field converged on the same design. ByteDance's Seedance line moved to a unified architecture that co-processes visual and audio signals in one latent space rather than syncing them after the fact, with multi-track output for background music, ambience, and character voiceover. Alibaba's Wan 2.5 added audio-visual synced generation with matching voices, effects, and music. Kling shifted toward an audio-visual model with native lip-sync in both English and Chinese. And xAI's Grok Imagine generates dialogue, effects, and music in the same pass as its short clips. In roughly a year, native audio went from one lab's showcase to a table-stakes feature across the field.

The limits are structural, not teething problems, and knowing them tells you exactly where standalone audio tools still win. Video models produce short clips, roughly ten to fifteen seconds, with audio that is incidental to the shot, not a four-minute song or a scored cue built to picture. They offer almost no fine control: you cannot open the mix, isolate a stem, retime a hit, or direct a performance the way a real sound-design workflow can. Their audio carries murky provenance, generated from training data whose rights are unclear, which makes it risky for work that has to be commercially cleared. And they do nothing at all for audio-only products, the podcast, the audiobook, the voice agent, where there is no picture to ride along with.

Check yourself0 / 5

Q01

Why is native audio a structural shift rather than just a new feature?

Q02

How reliable is native audio in practice, and where is it weakest?

Q03

What structural limits keep standalone audio tools relevant?

Q04

What is the likely equilibrium between native audio and standalone tools?

Q05

What does it mean that native audio went from showcase to table stakes in a year?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.