For most of the video era the output was silent. You generated a clipCLIPA text encoder developed by OpenAI in 2021 that learns to align text and images in a shared embedding space. Foundation of most text-to-image models from 2022 onward., then found music, foleyFoleyThe everyday sound effects of a scene, footsteps, cloth, impacts, now increasingly generated from the video itself., and dialogue elsewhere and married them in post. In 2026 that stopped being the default.
Frontier video models now generate synchronized audio in the same pass that makes the picture. Google's Veo introduced native dialogue, sound effects, and ambient audio in one pass in 2025. Kuaishou's Kling 3.0, in early 2026, shipped a unified multimodalMultimodalHandling more than one kind of data, for example text and images together, in a single model. architecture it calls Omni Native AudioNative audioSound generated jointly with video in a single pass, so footsteps, dialogue, and ambience line up with the picture, rather than being added afterward., generating multilingual dialogue and ambient sound locked to the video. Lightricks open-sourced LTX-2 in January 2026, a production model with native audio and lip-sync at up to 4K and 50 frames per second, released under genuinely open weightsOpen weightsModels whose trained parameters are publicly downloadable, allowing anyone to run them locally or fine-tune them. Contrasts with closed/API-only models., free below ten million dollars of revenue.
The architecture is the point. These models learn audio and video jointly, so the footstep lands on the frame the foot hits the ground and the lips move on the right phonemes. That kind of synchronization is very hard to reconstruct in post, which is why single-pass generation is a real capability rather than a convenience.
A parallel open-source track handles the reverse direction, video to audio. Foley-Omni, released in mid-2026 under an MIT license, takes a silent clip and generates a full soundtrack of dialogue, effects, and music in one pass. Tencent's HunyuanVideo-Foley is another strong open option. Sound is no longer something only the frontier APIs can give you.
The operator takeaway is that a video model is now an audio model. When you compare Veo, Kling, Seedance, or the open Wan and LTX lineages, native audio and lip-sync belong in the comparison table, not in a footnote. A model that only makes silent video is doing half the job.
Check your understanding
pass: 5 of 7
Answer at least 5 of 7 correctly to unlock the next chapter.
1. What fundamentally changed about frontier video model output in 2026?
2. What is the name of the unified multimodal audio architecture Kling 3.0 shipped in early 2026?
3. Why is single-pass audio-video generation a real capability rather than just a convenience?
4. What does Lightricks' LTX-2, open-sourced in January 2026, provide?
5. What does Foley-Omni, released mid-2026 under an MIT license, do?
6. What is the stated operator takeaway about evaluating video models in 2026?
7. How is the overall status of native synchronized audio in 2026 summarized?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.