Native audio in video, and the pressure it creates
Chapter 58
3 min read
Reviewed v78 · August 2026
Native audio in video, and the pressure it creates
The single biggest structural force in audio right now is not a new audio model at all. It is that the video models have started generating their own sound. Until 2025, AI video was effectively silent film: you rendered the picture, then found or made audio separately and tried to line it up. Native audioNative audioSound generated jointly with video in a single pass, so footsteps, dialogue, and ambience line up with the picture, rather than being added afterward. changes the unit of production. The model emits picture and sound together in one pass, so footsteps land on the frame the foot touches, dialogue tracks the lips, and room tone matches the room. Because the two signals are generated jointly rather than stitched together, synchronization is built into the render rather than bolted on afterward.
Google's Veo 3 was the first major model to ship this, announced at Google I/O in May 2025, generating dialogue, sound effects, and ambience natively alongside the video. OpenAI's Sora 2 followed on September 30, 2025, producing synchronized dialogue, effects, and soundscapes inside the same generation pipeline. Through 2026 the rest of the field converged on the same design. ByteDance's Seedance line moved to a unified architecture that co-processes visual and audio signals in one latent spaceLatent spaceThe compressed numeric representation an image is turned into before diffusion. Working here instead of on raw pixels is what makes modern image models fast enough to run. rather than syncing them after the fact, with multi-track output for background music, ambience, and character voiceover. Alibaba's Wan 2.5 added audio-visual synced generation with matching voices, effects, and music. Kling shifted toward an audio-visual model with native lip-sync in both English and Chinese. And xAI's Grok Imagine generates dialogue, effects, and music in the same pass as its short clips. In roughly a year, native audio went from one lab's showcase to a table-stakes feature across the field.
The limits are structural, not teething problems, and knowing them tells you exactly where standalone audio tools still win. Video models produce short clips, roughly ten to fifteen seconds, with audio that is incidental to the shot, not a four-minute song or a scored cue built to picture. They offer almost no fine control: you cannot open the mix, isolate a stem, retime a hit, or direct a performance the way a real sound-design workflow can. Their audio carries murky provenanceProvenanceA verifiable record of how a piece of media was made and by whom, increasingly required so that synthetic content can be traced., generated from training data whose rights are unclear, which makes it risky for work that has to be commercially cleared. And they do nothing at all for audio-only products, the podcast, the audiobook, the voice agent, where there is no picture to ride along with.
Check yourself0 / 5
Q01
Why is native audio a structural shift rather than just a new feature?
Q02
How reliable is native audio in practice, and where is it weakest?
Q03
What structural limits keep standalone audio tools relevant?
Q04
What is the likely equilibrium between native audio and standalone tools?
Q05
What does it mean that native audio went from showcase to table stakes in a year?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.