Realistic AI human movement fails for a reason most guides skip. Video models predict likely frames, and the likeliest version of a human action is the easy one. Effort gets averaged out, so bodies move without weight. The fix is structural, not verbal: supply the motion as a reference, put the load in the first frame, and cut the take short.
You have a six-second clip of a model lifting your bottle off a shelf. Nothing is distorted. The hand is intact, the face holds, the shelf stays a shelf. It still reads wrong, and nobody on the team can say why. The arm never loads. The bottle arrives in her hand at the weight of an empty one.
That is a different failure from the ones most AI video guides cover, and it takes a different fix. This guide separates motion realism from visual distortion, explains why your eye catches it so fast, and gives you four levers in the order that costs least.
Key Takeaways
- Motion realism is not visual quality. A clip can be clean at every frame and still move wrong. The two failures live on separate axes and respond to separate fixes.
- The tell is missing effort. Runway names “success bias” as a known limitation of Gen-4.5: actions disproportionately succeed, so a poorly aimed kick still scores (runway.com, July 2026). Effort disappears with it.
- Your eye is specialised for this. Observers read a person, their action, and their emotional state from 12 moving dots alone. A single frozen frame of the same dots reads as nothing (Johansson, 1973).
- Fixing the still does nothing for the movement. Motion information lives between frames. Repairing the source image solves distortion, not weightlessness.
- Transfer the motion instead of describing it. Kling 2.6’s Motion Control replicates specific movements from an uploaded video (ir.kuaishou.com, December 2025). A reference clip carries physics a prompt cannot.
- Shorten the take before you rewrite the prompt. Error accumulates along a sequence. Three two-second shots hold better than one six-second shot, and cost the same.
What makes AI human movement look fake
AI human movement looks fake when the body moves without weight. Joints travel smoothly, feet skim the floor, an arm lifts a heavy object at the speed of an empty hand. Nothing in the frame is broken, so the clip passes a visual check and fails a human one. The cause is that video models predict probable pixels, and probable motion is motion without strain.
That distinction matters commercially, because the standard advice stack acts on the wrong layer. Write a richer prompt, upgrade to the premium model, add more adjectives about realism. All three improve how the frame looks. None of them change the physics the model is extrapolating, because the model was never failing to understand your words.
Why your eye catches human motion before anything else
Human motion perception is a specialised system, and it is unusually hard to fool. The classic demonstration is the point-light walker, introduced by Johansson in 1973: 12 lights attached to the joints of an otherwise invisible person. Observers read a walking human immediately, and from those dots alone they can identify the action, the gender, the emotional state, and sometimes the individual.
The part that matters for AI video is the control condition. Shown a single still photograph of the same 12 lights, observers extract nothing. No figure, no person, no read at all. The percept only appears once the dots move.
That tells you where the realism signal lives. It is not inside any frame. It is in the relationships between frames, in how a knee leads a foot and how a torso counterweights an arm. Every viewer of your ad has a lifetime of training on exactly that signal, which is why a clip with flawless skin and perfect lighting can still get flagged as off within a second.
It also explains a common review dynamic. Your designer sees a clean frame and approves it. Your customer sees six seconds of motion and bounces. They are reading different information from the same file.
Missing effort is the tell, not broken anatomy
Here is the mechanism, stated by a model vendor about its own model. Runway published three persistent limitations alongside Gen-4.5, a model it describes as moving objects with “realistic weight, momentum and force” (runway.com, July 2026):
- Causal reasoning. Effects sometimes precede causes, such as a door opening before the handle is pressed.
- Object permanence. Objects disappear or appear between frames, such as a cup vanishing after it is hidden.
- Success bias. Actions disproportionately succeed. Runway’s own example is a poorly aimed kick that still scores.
Success bias is the one that governs human movement, and it is the least discussed. A generative model resolves an action toward its completed, successful state, because completion is the likelier outcome in the training data. Struggle, hesitation, near-miss, and correction are all lower-probability paths.
Strip those out and you strip out effort. Effort is the visible cost of moving mass, and it is most of what makes movement look real. A person lifting something heavy braces first, moves slowly through the hardest part, and resets. A model that skips to the successful outcome renders the lift without the bracing, which is precisely the weightless quality people describe when they say a clip looks generic AI.
The practical read: stop asking the model to depict effort and start supplying it. Effort is structure, and structure comes from inputs.
Why fixing the still frame does not fix the movement
There is good advice in circulation that says repair the source image before you animate it. That advice is correct, and it is for a different problem. An image edit costs 5 credits against thousands for a video take, so catching a bad hand or a warped label in the still is the cheapest fix available. We cover that workflow in the four distortion modes, and the same logic drives why hands and faces break first. The still does set which camera moves are possible, which is a separate question handled in how to animate a photo with AI.
Motion realism does not respond to it. A perfect still contains zero biological motion information, as the point-light control condition shows. You can iterate the source frame 50 times and the clip will still move without weight, because you have been improving a layer the failure does not live on.
Use this split when you triage a clip:
| What you see | Which failure | Where the fix goes |
|---|---|---|
| Warping, morphing, duplicated background, hand through product | Distortion | The source still, or the take length |
| Face changes across the clip | Identity drift | A reference image, not a description |
| Clean frames, floaty or effortless movement | Motion realism | The motion input and the first frame |
| Action completes too easily, no strain visible | Success bias | Cut around the effort, or transfer real motion |
Two clips can look equally polished and need opposite work. Naming the mode first is what stops the blind re-roll.
Four fixes, in the order that costs least
Work down this list. Each step is cheaper than the one after it. All four are really one decision, which input mode the shot starts from, applied to motion.
1. Transfer the motion instead of describing it
The strongest lever is handing the model real movement. Kling 2.6 shipped Motion Control, which “enables users to replicate specific movements from uploaded videos or from the online motion library” (ir.kuaishou.com, December 2025). Seedance 2.0 accepts multi-reference input across images, video, and audio, and references motion rhythm among the elements it carries across (seed.bytedance.com, 2026).
A four-second phone clip of someone performing the action gives the model a weight profile no sentence can. This is the single highest-yield change on the list, and most teams never try it because the prompt box is the obvious surface.
2. Put the load in the first frame
Anchor the start of the shot on a body that is already under strain. Veo 3.1’s First and Last Frame lets you define the opening and closing composition and generate the transition between them (blog.google, 2026). Ingredients to Video accepts up to three asset images of a person, character, or product and holds that subject through the clip.
A first frame showing a braced stance, a bent knee, or a tensed forearm sets the physical state the model extrapolates from. Start on a neutral standing pose and it will extrapolate neutral.
3. Cut on the effort, not through it
Success bias is worst across the moment an action resolves. Do not ask a model to render the whole lift. Render the brace, cut, then render the object already in hand. The viewer supplies the middle, which is what film editing has always relied on.
This is also the fix for anything that should shatter, pour, or collide. Cutting around the event costs one extra short take and removes the failure entirely.
4. Shorten the take
Error compounds along a sequence, so the tail of a clip degrades before the head does. Three two-second shots hold motion better than one six-second shot at identical credit cost, because video is billed per second of output. You also get three chances to keep a good one instead of one all-or-nothing render.
Prompt refinement sits below all four. It helps framing, subject, and camera, and the seven layers of a realistic AI video prompt is the right reference once the structural work is done.
Which models give you a motion lever
The question is not which model looks best. It is which model lets you supply motion as an input. Verified from each provider, as of July 2026:
| Model | Motion lever it exposes | Source |
|---|---|---|
| Kling 2.6 Pro | Motion Control replicates specific movements from an uploaded video or a motion library | ir.kuaishou.com, Dec 2025 |
| Veo 3.1 | First and Last Frame composition control; Ingredients to Video holds a subject from up to 3 asset images | blog.google, 2026 |
| Seedance 2.0 | Multi-reference input across text, image, video, and audio; carries motion rhythm from the reference | seed.bytedance.com, 2026 |
| Runway Gen-4.5 | Weight, momentum, and force in object motion, with success bias named as a known limit | runway.com, July 2026 |
Read the fourth row as the honest one. Runway publishing its own failure modes is more useful to you than any marketing claim about realism, because it tells you exactly which shots to plan around.
DesignerBox includes all four, plus Veo 3.1 Fast for cheap iteration, on one subscription. The full catalog is 13 image and video models across six providers at the models page. Switching models mid-campaign does not mean a second bill or a second prompt syntax to learn.
What testing motion costs in credits
Video is by far the most expensive operation on the platform, priced per second of output rather than per generation. That is why the ordering above matters.
| Action | Credits |
|---|---|
| Generate or edit an image | 5 |
| Seedance Pro Fast, 720p, 5 seconds | 150 |
| Kling Standard, 720p, 5 seconds | 225 |
| Sora 2, 720p, 8 seconds | 1,600 |
| Veo 3 with audio, 8 seconds | 6,400 |
Plans run 112 credits free, 500 on Basic at $15 a month, 1,000 on Pro at $35, 2,500 on Premium at $75, and 8,000 on Ultra at $200. A single Veo 3 take with audio at eight seconds costs more than Premium’s entire monthly allocation, so a blind motion re-roll on the top model is a budget event, not an iteration.
Run motion tests on the cheapest model that exposes the lever you need, confirm the movement holds, then render the approved setup once at quality. Full breakdown in what AI video really costs, and current plan detail on the pricing page.
Where this still breaks
Four things resist every fix above, and you should plan shots around them rather than budget for them.
Sustained fine motor work. Tying, threading, fastening, buttoning. High structural detail in few pixels, held over many frames.
Two people making contact. Handshakes, passes, embraces. Two bodies plus an occluded contact point is the hardest case in the frame.
Anything that must fail. A drop, a stumble, a spill, a missed catch. Success bias points directly against you, and no prompt overrides it.
Long continuous performance. Character identity holds better than motion quality over duration. Past a few seconds, plan cuts. Anchoring a character to a reference image solves identity across those cuts, which is a separate job from motion.
None of this is a reason to skip AI video for people. It is a reason to write shot lists that avoid the four cases, which is what a competent director does with a real crew and a real budget anyway.
Start from a shot you can direct
Motion realism is a directing problem wearing a technical costume. Supply the movement, set the physical state in frame one, cut on the effort, and keep takes short. Those four decisions do more than any prompt rewrite or model upgrade.
DesignerBox runs every model above from one canvas, on one bill, with your product and your people as the source material rather than a text description. Build the campaign in Ad Studio and switch models per shot without switching tools.
Which models hold anatomy across a full clip, and which fail predictably, is catalogued in AI creative failure modes.
FAQ
Why does AI video movement look floaty or weightless?
Because generative models predict probable frames rather than simulating physics, and probable motion is motion without strain. Runway names success bias as a known limitation of Gen-4.5, where actions disproportionately succeed (runway.com, July 2026). When the struggle is averaged out, the visible cost of moving mass goes with it, and bodies read as weightless.
Can a better prompt fix unrealistic human motion?
Not on its own. Prompts control framing, subject, camera, and style, and they help those layers reliably. Motion realism comes from what the model has to extrapolate, so it responds to inputs: a reference video, a loaded first frame, a shorter take. Rewrite the prompt after the structural work, not instead of it.
Which AI video model handles human movement best?
The better question is which model lets you supply the motion. Kling 2.6 Pro exposes Motion Control for replicating movement from an uploaded video (ir.kuaishou.com, December 2025). Veo 3.1 exposes First and Last Frame plus reference assets (blog.google, 2026). Seedance 2.0 takes multi-reference input including video. A model with a motion input beats a model with a higher benchmark score.
Does a longer clip make motion worse?
Yes. Error accumulates along a sequence, so the end of a clip degrades before the beginning does. Three two-second shots hold movement better than one six-second shot, and cost the same because video is billed per second of output.
Is motion transfer the same as image-to-video?
No. Image-to-video animates a still and invents the movement. Motion transfer takes a separate reference video and applies its movement to your subject, so the weight profile comes from real footage rather than from prediction. Transfer is the stronger tool when the movement itself is the point.
Can I fix bad motion in post instead of re-generating?
Rarely. Colour, speed, and cuts are adjustable in post. Weight is not, because it is encoded in the relationships between frames. Trimming to the strongest second and cutting away is usually faster and cheaper than any correction attempt.
How do I test motion without burning a month of credits?
Test on the cheapest model that exposes the lever you need, at the shortest duration that shows the movement. A five-second Seedance Pro Fast take at 150 credits proves whether the motion holds. Only then render the approved setup on the model you plan to ship.
Sources
- Runway Gen-4.5’s three published limitations, including success bias, and its description of weight, momentum and force in object motion: (runway.com, July 2026)
- Kling 2.6 Motion Control, replicating specific movements from an uploaded video or the online motion library: (ir.kuaishou.com, December 2025)
- Veo 3.1 First and Last Frame control, and Ingredients to Video holding a subject from up to three asset images: (blog.google, 2026)
- Seedance 2.0 multi-reference input across text, image, video and audio, and motion rhythm carried from the reference: (seed.bytedance.com, 2026)
- The point-light walker experiment and the still-frame control condition: Johansson, 1973, and subsequent point-light display research
- DesignerBox pricing, credit costs, plan allocations and feature gating verified against live product configuration, July 2026
Model capabilities verified from Runway Research, Google, ByteDance Seed, and Kuaishou investor communications as of July 2026. Biological motion findings from Johansson (1973) and subsequent point-light display research. Credit costs from DesignerBox platform configuration. Individual results vary.