Skip to main content
Get started free

Realistic AI Lip Sync: Why Talking Video Looks Fake

Every fake tell in AI talking video traces to one choice: generating picture before audio. What the sync research shows, and the order that fixes it.

Realistic AI Lip Sync: Why Talking Video Looks Fake

AI lip sync looks fake when the mouth is accurate but the performance is not. Current models hit sync tolerances the human eye cannot resolve. What reads as fake is everything around the mouth: a face that holds one expression, eyes that blink on a timer, and a body that gestures without emphasis. Those come from generating the picture before the audio exists.

You have probably seen the result in your own feed. The mouth tracks the words correctly. The voice is clean. Something is still wrong, and nobody in the review call can name it, so the note comes back as “it feels off” and you re-roll the clip. Three generations later it still feels off, because the thing you keep changing is not the thing that is broken.

This covers what viewers really detect, the millisecond budget broadcast engineers have measured for sixty years, the four tells that give away a blind performance, and the generation order that removes them. It is about output quality. If you are still deciding whether your ad needs a lip sync pass at all, when a video ad actually needs AI lip sync covers that question and the consent rules attached to it.

Key Takeaways

Sync accuracy is not your problem. Broadcast standards put the detectability threshold at 45 milliseconds of audio lead to 125 milliseconds of lag (ITU-R BT.1359-1). Generated mouths land well inside that window. The fake read comes from somewhere else.

Vision overrides hearing in speech perception. In the original McGurk study, 98% of participants heard a syllable that was never played, because the mouth showed a different one (McGurk and MacDonald, Nature, 1976). Viewers are not watching the mouth for accuracy. They are reading it for intent.

Blink rate is a measurable tell. Speaking raises spontaneous blink rate to about 26 per minute, against 17 at rest and 4.5 while reading (Bentivoglio et al., Movement Disorders, 1997, n=150). An 8-second talking clip should show roughly three to four blinks. Two reads as a mask.

The cause is generation order. A model that renders the picture before the audio exists cannot place emphasis, breath, or expression against a line it has not heard. Every downstream fix patches a symptom.

Native-audio models remove the ordering problem by design. Veo 3.1 generates audio with video and it cannot be switched off (ai.google.dev, July 2026). Seedance 2.0 generates picture and synchronised audio in a single pass.

DesignerBox does not lip sync footage you already shot. It generates new video with native audio across 13 models. Both facts matter for picking the right tool, and the second is covered below.

Why does AI lip sync look fake when the mouth matches the audio?

Because viewers do not audit the sync frame by frame. They read a face for intent, and intent lives in timing, not in phoneme shapes. A mouth can form every sound correctly while the face behind it communicates nothing that matches the words.

The perceptual research here is old and settled. When Harry McGurk and John MacDonald dubbed the audio of one syllable onto the video of another, 98% of participants reported hearing a third syllable that was present in neither track (McGurk and MacDonald, Nature, 1976). The paper has been cited more than 4,800 times. What it established is that vision is not a secondary channel in speech perception. It actively rewrites what people believe they heard.

That has a direct consequence for generated video. Your viewer’s brain is fusing the face and the voice into one judgment. When the face carries no emphasis and the voice does, the fusion fails, and the failure surfaces as a vague sense of wrongness rather than a specific complaint. This is why creative feedback on AI talking video is so often unhelpful. The viewer is detecting something real and has no vocabulary for it.

How far off can lip sync be before a viewer notices?

Further than you would guess, and the tolerance is asymmetric. Sound travels slower than light, so human perception is built to expect audio slightly after vision, and it forgives lag far more than lead.

Broadcast engineering has measured this precisely:

StandardAudio may lead byAudio may lag byWhat it measures
ITU-R BT.1359-145 ms125 msThreshold of detectability
EBU R3740 ms60 msEnd-to-end programme tolerance
ATSC IS-19115 ms45 msBroadcast operating limit
Film practice22 ms22 msTheatrical projection

At 24 frames per second, one frame is about 42 milliseconds. The ITU detectability window spans roughly three frames. Generated output from current models sits inside it comfortably.

Which settles the diagnostic question. If your clip reads as fake, measuring the sync is wasted effort. The error is not in the millisecond alignment between mouth and waveform. It is in what the rest of the face was doing while the mouth moved.

The cause sits upstream of the mouth

Here is the decision that produces the problem, and it is made before anyone writes a prompt.

Two pipelines exist for talking video. In the first, a model generates the picture, then a second pass animates the mouth to a supplied audio track. In the second, one model generates picture and audio together in a single pass.

The first pipeline has a structural gap. At the moment the video is generated, the line does not exist yet. The model chooses a facial performance, a blink cadence, a gesture rhythm, and a breath pattern with no information about what is being said, where the stress falls, or how long the sentence runs. The lip sync pass then arrives and corrects the mouth. It corrects only the mouth. Everything else in the frame was decided blind and stays blind.

Model strength is beside the point here. A stronger model generating blind produces a more convincing face doing the wrong thing. Adding prompt detail about expression does not close the gap either, since a prompt describes an average intention across the whole clip while emphasis is a per-word event.

That reframes the fix. The tells below are not five separate problems needing five separate solutions. They are four readings of one missing input.

Four tells, and the one decision behind all of them

Each of these is what a blind performance looks like from the outside. Check a suspect clip against them in order.

Emphasis lands flat. Real speech puts weight on one or two words per sentence, and the face moves with it: a brow lift, a head dip, a widening on the stressed syllable. A performance generated before the line existed distributes movement evenly, because it had no way to know which word mattered. Watch a clip with the sound off. If you cannot tell where the sentence peaked, the viewer cannot either.

The blink clock runs at rest. Spontaneous blink rate is task-dependent and well quantified: about 17 blinks per minute at rest, rising to 26 during conversation, and dropping to 4.5 while reading (Bentivoglio et al., Movement Disorders, 1997). Speaking is the high-rate state. An 8-second clip of someone talking should carry roughly three to four blinks. Generated talking heads routinely deliver two, evenly spaced, which reads as a mask rather than a face. Count them. It takes one pass.

Breath is missing. Speech is a physical act with preparation. A long sentence starts with an inhale, the shoulders and chest move, and a pause before an important line is loaded rather than empty. Generated performances tend to have pauses that are simply gaps. This is the tell viewers are least able to articulate and the one that most reliably makes a clip feel synthetic.

Expression is averaged across the clip. A line that opens sceptical and closes warm needs the face to travel. A model working from a prompt applies one emotional setting to the whole generation. The result is a face that is consistently pleasant across a sentence that was not consistently pleasant.

Two related failures sit outside this list. A face drifting into a different face across a multi-clip sequence is a consistency problem, covered in how to keep characters consistent in AI video. Body motion that reads as weightless is a physics problem rather than a speech one, and why realistic AI human movement fails covers that separately.

How to generate talking video audio-first

The correction is an ordering change, not a tool purchase. Give the model the line before it commits to a performance.

  1. Write and lock the script first. Not a topic, the actual words, punctuated as they will be spoken. The script is now an input to the picture, so it cannot be finalised afterwards.
  2. Mark the stress. Pick the one or two words per sentence that carry the point. This is the information the model is otherwise missing, and it is the difference between a face that agrees with the voice and one that runs beside it.
  3. Generate the audio, or let the model generate both. Either produce the voice track before any picture exists, or use a model that generates speech and video jointly so the two are never decided separately. The cost and read-length trade-offs between those two options are compared in AI voiceover for UGC ads.
  4. Write the prompt around the delivery, not the appearance. Describe the breath before the second line, the brow on the stressed word, where the eyes go on the pause. Appearance detail is cheap and the model handles it. Delivery detail is what is scarce. Our guide to writing realistic AI video prompts covers the layer structure this fits into.
  5. Cut the line to the model’s duration before you generate. A script that overruns forces the delivery to compress, and compressed delivery is the first thing to read as synthetic. Count the words against the clip length rather than trimming after.
  6. Review with the sound off, then with picture off. Two passes, two different faults. Silent review exposes flat emphasis and blink rate. Audio-only review exposes a read that was never going to sit on a human face.

Step six is the one teams skip and the one that catches the most. A clip that survives both passes separately will survive them together.

Which models generate picture and audio in one pass

Joint generation is the architectural version of audio-first. The model is never in a position to decide the performance blind, because the speech is part of the same generation.

ModelNative audioDurationSource
Veo 3.1Always on, cannot be disabled4, 6 or 8 secondsai.google.dev, July 2026
Seedance 2.0Yes, generated in a single passUp to 15 secondsseed.bytedance.com, July 2026
Kling 2.6 ProYes, at 1080p professional mode5 or 10 secondsfal.ai, July 2026
Runway Gen-4.5Not documented for gen4.5Not publisheddocs.dev.runwayml.com, July 2026

Three specifics worth knowing before you plan around them.

Veo 3.1’s audio is not optional. Google’s documentation lists native audio as always on, with durations of 4, 6 or 8 seconds, and 8 required for 1080p and 4K output (ai.google.dev, July 2026). The parameter that governs whether people appear is also regionally constrained: in the EU, UK, Switzerland and MENA, only the adult setting is permitted (ai.google.dev, July 2026).

Sora 2 Pro will not take a photograph of a real person. OpenAI’s API documentation states that real people including public figures cannot be generated, and that input images with human faces are rejected (developers.openai.com, July 2026). It is a strong model for talking video with a synthetic character and the wrong tool for your founder.

Kling 2.6 Pro’s audio requires 1080p professional mode, so a sound-on clip costs more per second than the headline rate implies.

You can browse the full catalogue at designerbox.ai/models, or go straight to Veo 3.1 and Seedance 2.0.

What audio-first costs in credits

Ordering the work correctly costs nothing. Generating video does, and it is the most expensive operation on the platform, priced per second of output.

Current DesignerBox credit costs for representative clips:

ClipCredits
Seedance Pro Fast, 720p, 5 seconds150
Kling Standard, 720p, 5 seconds225
Sora 2, 720p, 8 seconds1,600
Veo 3 with audio, 8 seconds6,400

A single Veo 3 clip with audio at 8 seconds runs 6,400 credits, which exceeds the Premium tier’s entire monthly allocation of 2,500. That number is the reason audio-first matters commercially rather than only aesthetically. Every blind generation you throw away at 6,400 credits is a month of Premium spent on a clip nobody will ship.

The plan ladder runs Free at 112 credits, Basic at $15 for 500, Pro at $35 for 1,000, Premium at $75 for 2,500, and Ultra at $200 for 8,000. AI video requires Premium or higher. Full pricing sits at designerbox.ai/pricing, and our breakdown of AI video generation cost runs the per-model arithmetic against real shot counts.

When you cannot go audio-first

Three cases, and in all three the ordering fix does not apply.

The footage already exists. If the ad was shot last quarter and the offer changed, no generator helps you. That is a dubbing or lip sync job on the original file, and the performance you are correcting was captured by a human, so the four tells above do not arise.

A real, identifiable person appears. Sora 2 Pro blocks this outright and Veo 3.1 constrains it regionally. Generating a real person also carries consent and disclosure obligations that sit outside output quality entirely.

The read runs long. Every native-audio model currently shipping caps well under a 30-second presenter script. Veo 3.1 stops at 8 seconds and Seedance 2.0 at 15. A long read needs either a dedicated avatar tool or a multi-clip sequence, and a multi-clip sequence brings the consistency problem back.

Where DesignerBox fits. DesignerBox generates new video with native audio across 13 image and video models on one subscription. It does not re-sync mouth movements in footage you already shot. If that is your job, a dedicated dubbing tool is the right buy and this article’s ordering advice does not apply to you. What DesignerBox covers is the other half: video built from scratch, with speech, from your actual product photo, with Ad Studio assembling the ad set around it.

The AI talking avatar tool runs the sync step free, and DesignerBox vs HeyGen compares it against a dedicated avatar platform. HeyGen alternatives compares six tools that handle the presenter step.

FAQ

Why does my AI lip sync look fake even though the mouth matches?

Because the fake read is not coming from the mouth. Current models land inside the 45 to 125 millisecond detectability window that broadcast standards define (ITU-R BT.1359-1), which the eye cannot resolve. What viewers detect is flat emphasis, a resting blink rate during speech, absent breath, and a single averaged expression across a line that should have travelled.

How many milliseconds of lip sync error can people detect?

The threshold of detectability runs from 45 milliseconds of audio leading video to 125 milliseconds of audio lagging it (ITU-R BT.1359-1). Tolerance is asymmetric because sound naturally arrives after vision. Stricter operating standards exist for broadcast: ATSC IS-191 sets 15 milliseconds lead and 45 lag, and EBU R37 sets 40 and 60.

Should I generate the audio before or after the video?

Before, or jointly. A model generating picture first has no access to the line, so it places emphasis, blinks, breath and expression with no information about what is being said. Generating audio first, or using a model that produces both in one pass, removes that gap rather than patching it downstream.

Roughly three to four. Spontaneous blink rate during conversation averages 26 per minute against 17 at rest (Bentivoglio et al., Movement Disorders, 1997), which works out to about 0.43 blinks per second while speaking. Two blinks in 8 seconds is a resting cadence on a talking face, and it reads as a mask.

Does DesignerBox do lip sync?

No. DesignerBox generates new video with native audio through models including Veo 3.1, Seedance 2.0 and Kling 2.6 Pro. It does not re-sync mouth movements in footage you already shot. For that job, a dedicated dubbing tool is the correct buy.

Can I use Veo 3.1 for a talking video?

Yes, within its limits. Audio generation is always on and cannot be disabled, and durations are 4, 6 or 8 seconds, with 8 required for 1080p and 4K (ai.google.dev, July 2026). Person generation is regionally restricted to the adult setting in the EU, UK, Switzerland and MENA.

Does a longer script make AI lip sync worse?

Yes, indirectly. Every native-audio model caps at a short duration, so an overlong script forces the delivery to compress, and compressed speech is one of the first things viewers read as synthetic. Cut the line to the clip length before generating rather than trimming the output afterwards.

Where to go next

If your clip reads as fake, run the four tells with the sound off before you re-roll. The fix is almost never a different model.

Related reading: choosing between dubbing and generation, how to make a photo talk with AI, why AI video hands and faces break, the seven layers of a video prompt, and what AI video generation costs.

Sources

  • Audio-to-video synchronisation tolerances: ITU-R BT.1359-1, EBU R37 and ATSC IS-191, accessed July 2026
  • Cross-modal speech perception, the basis for why a mismatched mouth reads as wrong: McGurk and MacDonald, Nature, 1976
  • Spontaneous blink-rate data used for the blink tell: Bentivoglio et al., Movement Disorders, 1997, n=150
  • Veo 3.1 always-on audio generation, 4, 6 and 8 second durations, and the regional restriction on person generation: (ai.google.dev, July 2026)
  • Sora 2 Pro synced audio and clip lengths: (developers.openai.com, July 2026)
  • Seedance 2.0 dual-channel audio including character voiceovers, and its multi-reference input: (seed.bytedance.com, July 2026)
  • Runway Gen-4.5 spec table, which contains no audio row: (docs.dev.runwayml.com, July 2026)
  • DesignerBox credit costs and plan allocations verified against live product configuration, July 2026

Synchronisation tolerances verified against ITU-R BT.1359-1, EBU R37 and ATSC IS-191 as of July 2026. Perceptual findings from McGurk and MacDonald (Nature, 1976) and Bentivoglio et al. (Movement Disorders, 1997). Model capabilities verified against provider documentation from Google (ai.google.dev), OpenAI (developers.openai.com), ByteDance (seed.bytedance.com) and Runway (docs.dev.runwayml.com) as of July 2026. DesignerBox credit costs and plan allocations current as of July 2026. Model specs in this category change monthly. Verify before committing a campaign. Individual results vary.

Vytas

Founder at DesignerBox

Vytas is a founder at DesignerBox, from the team behind LoadFocus, FocusBox and PostNext. He writes about turning one product photo into a full campaign, and the pipelines that keep every asset on brand.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Reshoot the clip, not the campaign

Veo 3.1, Sora 2 Pro, Kling 2.6 Pro, Seedance 2.0 and Runway Gen-4.5 are built in. When one model warps a hand, rerun the same shot on another without a second subscription.

Start free

Upload one product photo. Ship the whole campaign, without a photoshoot.