AI video breaks hands and faces because both pack more structural detail into fewer pixels than anything else in the frame, and the model has to rebuild that detail on every frame with no memory of the last one. Static hands are largely solved in 2026. Hands touching an object are not. The reliable fix is to start from a still where the anatomy is already correct.
You have seen the clip. Six seconds of someone lifting your bottle off a counter. The lighting is right, the motion is right, and at around frame 40 the thumb passes through the glass. Nobody on the team can unsee it, the clip is dead, and you have already spent the credits.
The instinct is to write a longer prompt. Add “five fingers”, add “anatomically correct hands”. That almost never works. This guide covers what breaks and what does not in 2026, the five tells to check before a clip ships, the one move that removes most of the failure, and what it costs to catch the problem in a still instead of a finished take.
Key Takeaways
- The failure moved. A hand at rest renders cleanly on current models. A hand gripping, rotating, or handing over an object is where distortion still lives, and that is the exact shot product video needs.
- Contact is a physics problem, not a drawing problem. On HanDyVQA, a June 2026 benchmark for hand-object interaction, the strongest model scored 72.6% against a human baseline of 96.6% (arxiv.org, June 2026).
- More prompt detail about anatomy makes it worse, not better. Describing fingers spends attention on the region that is already unstable. Describe the action and the camera instead.
- Fix the still, not the video. An image edit is 5 credits. A Veo 3 clip with audio at 8 seconds is 6,400. Repair the frame first, then animate the approved frame.
- Every video model in the catalogue accepts a reference or a start frame. That input is what turns generation from invention into preservation.
- Some shots still will not hold. Two people sharing a frame, a re-grip mid-take, and fine manipulation like fastening a clasp remain unreliable. Block around them.
Why hands and faces break more than anything else in the frame
Three things stack up on those two regions and nowhere else.
Detail density per pixel. A hand carries 27 bones and a joint structure that reads wrong the instant one angle is off, and in a product shot it usually occupies a small share of the frame. The model resolves high-frequency structure inside a low-pixel budget. Faces have the same problem with a harder audience: human vision is tuned to faces, so an error in eye spacing or jaw width registers before a viewer can name it.
No memory between frames. Video models resolve structure per frame. Small errors do not cancel across a sequence, they flicker. A finger that is slightly too long on frame 12 and correct on frame 13 produces a shimmer no still would ever show, which is why a clip can fail on motion while every frame looks fine paused.
Contact is the hard part. When a hand reaches a product, the model holds wrist rotation, finger placement, depth, occlusion, and the point where skin meets surface at once, and keeps them agreeing frame to frame. Research through 2026 names the same weak spots: spatial relationships between objects, use of temporal information for motion, and part-level geometry (arxiv.org, June 2026).
The failure moved, and most guides have not caught up
Guidance written in 2023 said AI could not draw hands. That is no longer the useful statement. Ask a current model for a hand resting on a table and you will usually get a hand. What still fails is the hand doing something.
The measured gap is in interaction, not anatomy. HanDyVQA, published June 2026, tests fine-grained understanding of hand-object interaction dynamics across 11.1K question pairs. The best model tested reached 72.6% accuracy. Professional human annotators averaged 96.6% (arxiv.org, June 2026).
That gap is where your clip falls, and it matters commercially because the interaction shot is the shot brands want. Picking the product up. Turning the label to camera. Applying the cream. Handing the box over. A talking-head clip with hands at rest will hold. A six-second unboxing will fight you.
Faces follow the same pattern. A single face holds inside a short take. A face across a cut, or across two clips, drifts, which is covered in our guide to keeping a character consistent across clips. Contact and anatomy is one of four distortion modes, and the fixes do not transfer between them. Diagnosing which mode you have covers the other three.
The five tells to check before a clip ships
Watch the clip at quarter speed once, then check these in order. Four of the five only appear in motion.
| Tell | What to look for | Why it happens |
|---|---|---|
| Finger count drift | Count fingers on the frame where the hand first touches the object, then again 10 frames later | Contact frames carry the most occlusion, so the model has the least structure to work from |
| Occlusion bleed | A finger passing through the product instead of behind or in front of it | Depth ordering is resolved per frame and can flip |
| Grip that does not track | The object moves but the hand shape stays fixed, or the reverse | Motion and contact are solved separately and can drift apart |
| Face micro-morph | Pause on the first and last frame, then flip between them. Jaw width, eye spacing, hairline | No frame-to-frame identity anchor within the take |
| Late-clip decay | Teeth, eyes, and finger edges in the final second | Error accumulates along the sequence, so the tail degrades first |
Fail on tells one to three and the problem is the source frame or the blocking. Fail on four and five and the take is too long, or the subject needs a reference anchor.
Fix the frame before you animate it
This is the move that removes most of the failure, and it is not a feature, it is an order of operations.
Text-to-video asks the model to do two jobs: invent the anatomy, then keep it stable through motion. Image-to-video asks for one: keep the given anatomy stable. Take away the invention step and you take away most of the places it goes wrong. That choice, and what it costs you elsewhere, is covered in choosing an AI video input mode.
Supplying a frame narrows the job without eliminating it. Any motion that turns a hand or a face reintroduces invention, because the model has to author the side your photo never showed. How to animate a photo with AI ranks motion by how much of that it forces.
- Start from a real photo where you can. Your actual product in a real hand is a source no model has to guess at. DesignerBox builds every asset from your product photo for this reason, and it is why nothing comes out looking generic AI.
- Or generate the still first, at 5 credits. Image generation and image editing are 5 credits each. Iterate the frame twenty times and you have spent 100 credits, less than one 5-second draft clip.
- Repair the still, not the video. Brush out a broken finger or a bad contact edge with the Retouch app at 5 credits per repair. For the face, Skin Retouch evens skin and clears blemishes while keeping texture, also 5 credits.
- Feed the approved still as the start frame. Every video model in the catalogue takes an image input. That frame becomes the anatomy the model preserves rather than the anatomy it invents.
The same sequencing is what makes one source image carry a whole ad, which we break down shot by shot in turning one product photo into a video ad.
What to change in the prompt, and what to change in the shot
Prompt changes help less than people expect. Blocking changes help more.
Stop describing the anatomy. Adding “five perfect fingers” spends prompt weight on a problem the model cannot solve by being told. Google’s own prompting formula for Veo 3.1 is cinematography, subject, action, context, style and ambiance, with camera language called out as the most powerful element (cloud.google.com, July 2026). Anatomy is not in that list.
Write exclusions as positive descriptions. Google recommends stating what you want excluded as part of a positive description of the scene rather than as a bare negative (cloud.google.com, July 2026). Applied to hands: “the bottle held low, out of frame below the chest” beats “no visible hands”.
Then change the shot, which is where the real gain is:
- One grip, held. A hand that takes hold once and keeps that hold survives. A re-grip mid-take almost never does.
- Keep hands fully in or fully out. A hand crossing the frame edge forces the model to invent and discard structure repeatedly.
- Slow the motion. Fast hand movement compounds per-frame error. A slow rotation reads as premium anyway.
- Shorten the take. Decay accumulates along the sequence. Two 4-second clips cut together beat one 8-second clip that falls apart at second seven.
- Establish contact before the camera arrives. Start with the product already held, then move the camera. The model preserves an existing grip more reliably than it forms a new one.
For the full structure of a prompt that specifies camera, light, and motion rather than anatomy, see how to write realistic AI video prompts.
Which models accept what
Every video model in the DesignerBox catalogue takes some form of image input. The limits differ, and the limit decides how much anatomy you get to lock.
| Model | Reference input | Best used for |
|---|---|---|
| Veo 3.1 | Ingredients to Video: up to 3 asset images of a person, character, or product (deepmind.google, July 2026) | The hero take, where the face and the audio both have to land |
| Veo 3.1 Fast | Same reference input, tuned for drafts | Blocking the motion before you spend flagship credits |
| Sora 2 Pro | First-frame image input, clip durations of 4, 8, 12, 16 or 20 seconds (developers.openai.com, July 2026) | Scene-led shots with synced audio |
| Seedance 2.0 | Multi-image reference anchoring one subject | Longer takes that need one subject to hold |
| Kling 2.6 Pro | Elements: up to 4 reference images, plus first and last frame control (fal.ai, December 2025) | Chaining short clips so a grip carries across a cut |
| Runway Gen-4.5 | References plus image-to-video from any still (runwayml.com, December 2025) | Animating an approved frame with minimal drift |
Two practical notes. First and last frame control, available on Kling 2.6 Pro and as interpolation on Veo 3.1, is underused here: you supply both ends of the motion, so the model interpolates between two frames you have already checked instead of inventing the destination. Second, a draft pass on a fast model is the cheapest diagnostic you have. If the hand fails on Veo 3.1 Fast, it will fail on the flagship, for a fraction of the credits. How the rest of the catalogue prices against that is in six video models compared on cost per second, clip length and audio.
For a wider read on which model suits which job, see our comparison of image-to-video models for ecommerce.
What it costs to catch it early instead of late
This is the argument that settles the workflow. Fixing anatomy in a still and fixing it in a video are not the same purchase.
| Operation | Credits |
|---|---|
| Generate an image | 5 |
| Edit or retouch an image | 5 |
| Nine-image avatar set | 25 |
| Seedance Pro Fast, 720p, 5 seconds | 150 |
| Kling Standard, 720p, 5 seconds | 225 |
| Sora 2, 720p, 8 seconds | 1,600 |
| Veo 3 with audio, 8 seconds | 6,400 |
Read the top row against the bottom one. A still repair is 5 credits. One 8-second Veo 3 take with audio is 6,400, more than two months of the Premium plan’s 2,500 monthly allocation on a single clip. Find the broken thumb after that render and the fix is a full re-render, not a touch-up.
Run it the other way and the arithmetic stops being close. Generate the frame, check the hands, repair what is wrong, approve it, then animate once. Video is by far the most expensive operation in the product, so the discipline is to spend the cheap credits on certainty and the expensive ones exactly once. Plan allocations sit on the pricing page, and the per-clip math is in how many clips your plan buys.
What still fails
Being honest about the ceiling saves you the credits you would spend finding it. Note that a clip can clear every check here and still move wrong, which is a separate failure covered in realistic AI human movement.
- Two people in one frame. Anchoring one subject works well. Two recurring subjects sharing a shot still blend traits, and hands from different people in contact is the worst version of it. Framing above the contact point is the cheapest fix, covered with the rest of the method in two characters in one AI image.
- Fine manipulation. Fastening a clasp, threading a lace, unscrewing a small cap, pouring precisely. Anything where the value of the shot is the dexterity itself.
- Re-grips. Picking up, setting down, and picking up again inside one take.
- Extreme close-ups on fingers. More pixels on the hardest region removes the forgiveness that scale was giving you.
- Long takes. Error accumulates. Where a shot has to run long, chain shorter clips with first and last frame control rather than asking for one continuous generation.
None of these are permanent. All of them are true in July 2026. When a shot lands in this list, the answer is a real photograph for that frame, or a blocking change that moves the interaction outside it. Neither is a compromise a viewer will notice, and both are cheaper than five failed renders.
Which failures are worth a re-render and which are fixable in post is catalogued in AI creative failure modes.
FAQ
Why do AI videos still show the wrong number of fingers?
Because a video model resolves structure on every frame independently, and the frames where a hand touches or overlaps an object carry the most occlusion. Less visible structure means more guessing, and the guess can change between one frame and the next. The count is usually correct in the frames where the hand is open and unobstructed, and drifts at contact.
Does describing the hand in more detail fix it?
No, and it often makes the clip worse. Adding “five fingers” or “anatomically correct hands” spends prompt attention on the region that is already unstable, without giving the model any new structure to work from. Google’s Veo 3.1 formula puts the weight on cinematography, subject, action, context, and style (cloud.google.com, July 2026). Give it a start frame instead of an adjective.
Which video model handles hands and faces best?
There is no single answer that survives a quarter, because the models update monthly. The reliable rule is structural: whichever model you use, give it an image input rather than a text description, and use first and last frame control where the shot allows. That changes the job from invention to preservation on every model in the catalogue.
Can I fix a bad hand without regenerating the whole clip?
Not inside the video. Video edits are not localised the way image edits are, so a broken frame means a re-render. That is the whole argument for repairing the source still first, where a fix is 5 credits and a re-render is not.
Why does a face look right at the start of a clip and wrong at the end?
Error accumulates along the sequence, so the tail of a take degrades before the head does. Shorten the clip, anchor the subject with a character reference image built to hold, and check the first and last frame against each other before you approve. Drift across separate clips is a related but separate problem, with a different fix.
Is a real product photo better than a generated one as the start frame?
Yes, when you have one. A real photo of your product in a real hand gives the model correct geometry, correct contact, and your actual product, with nothing to invent. A generated still is the fallback, and at 5 credits per iteration it is a cheap one. Either way, the frame gets checked before it gets animated.
Do two people in one shot make hand and face errors worse?
Yes. Anchoring a single subject to a reference works well across current models. Two recurring subjects in one frame still blend features, and interacting hands from two people is the hardest version of the contact problem. Generate them in separate passes where the edit allows, or block the shot so only one person’s hands are in frame.
Where to go next
Faces that break while speaking are a different fault with a different cause. Why AI talking video reads as fake covers blink rate, emphasis and the sync research.
Turn one product photo into the full campaign, on brand, in one workspace. See what the Ad Studio does.
Sources
- Hand-object interaction accuracy of 72.6% against a 96.6% human baseline, and the named weak spots in spatial relationships, temporal information and part-level geometry: HanDyVQA (arXiv:2512.00885, June 2026)
- Veo 3.1 prompting formula, camera language as the most powerful element, and the guidance to write exclusions as positive description: (cloud.google.com, July 2026)
- Veo 3.1 Ingredients to Video, up to 3 asset images of a person, character or product: (deepmind.google, July 2026)
- Sora 2 Pro first-frame image input and 4, 8, 12, 16 and 20 second clip durations: (developers.openai.com, July 2026)
- Kling 2.6 Pro Elements, up to 4 reference images, plus first and last frame control: (fal.ai, December 2025)
- Runway Gen-4.5 references and image-to-video from any still: (runwayml.com, December 2025)
- DesignerBox credit costs, per-model video credit examples and plan allocations verified against live product configuration, July 2026
Model capabilities and reference limits verified against Google DeepMind, OpenAI, fal.ai, and Runway documentation as of July 2026. Benchmark figures from HanDyVQA (arXiv:2512.00885, June 2026). Credit costs from DesignerBox plan configuration. Individual results vary.