Contents

62 / 153

Specialized Models: Editing, Lipsync, 3D, and Upscaling

Lipsync and audio-visual synchronization

Chapter 61

3 min read

Reviewed v78 · August 2026

Lipsync is the literal seam between the audio pillar and video. Generative video can produce a convincing person and generative audio can produce convincing speech, but a talking human is only believable when the mouth, jaw, and face move in time with the sound. Audio-driven lip and face animation is the family of models that solves exactly this: given a face, a photo, a clip, or generated footage, and an audio track, it drives the mouth and surrounding facial motion to match the speech. It is what turns a still portrait into a speaker and what lets a dubbed film actually look dubbed rather than dubbed-over.

The research lineage runs straight into today's products. Wav2Lip, from the 2020 paper A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild, remains the most widely deployed open-source lip-sync model and still powers hobbyist and research pipelines. Its authors went on to found Sync Labs, and the original repository openly points users to that commercial successor. Sync.so is now the production form of that research, with a flagship model that targets high resolution and high frame rate, works zero-shot on faces it has not seen, and supports dubbing across dozens of languages. The through-line is clear: Wav2Lip was the science, and Sync's models are what it became once rebuilt for resolution, speed, and language coverage.

The rest of the field has specialized. Hedra's Character-3 is an omnimodal talking-avatar model that takes a single portrait plus audio or text and generates a speaker with natural micro-expressions like blinks and small head movements. HeyGen concentrates on localization at scale, dubbing into many languages with re-synced lip movement, and its late-2025 precision translation engine added better occlusion handling and multi-speaker support. Runway's Act-Two takes a different route called performance transfer: instead of driving from audio, it maps a real performer's head, face, body, and hand motion, from a driving video, onto a reference character. And lipsync is increasingly a feature inside general video models rather than a separate tool, with Kling adding a lip-sync capability that re-animates a character's mouth to new audio or typed text while keeping the background intact. The boundary between video generation and audio-visual sync is dissolving.

The uses cluster into three areas. Dubbing and localization is the largest, making one performance play naturally in many languages. Avatars and virtual presenters come next, spanning marketing, training, and social content built from a single photo or a cloned likeness. And filmmaking dialogue is the emerging frontier, where directors want generated or de-aged characters to deliver lines convincingly.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. What does a lipsync model do?

  2. 2. Why is lipsync a notoriously difficult problem?

  3. 3. Which model is described as the dominant lipsync option in the current generation?

  4. 4. How is Sync Labs characterized?

  5. 5. Why might standalone lipsync tools become less critical over time?

  6. 6. Why are standalone lipsync tools still better than the big video models at the task for now?

  7. 7. Besides Sync Labs, what other options exist in the lipsync space?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.