AI video distorts because each frame is generated independently, with no memory of the one before it. The distortion arrives in four modes: identity drift, environment warp, contact and anatomy breaks, and physics failures. Each has a different fix. Prompt changes only help the first three, and a better-looking model does not fix the fourth.
You have a six-second clip of a model picking up your bottle. The face is right at second one and someone else’s by second five. The shelf behind her has grown a second shelf. The bottle passes through her thumb. You re-roll it, twice, and spend 12,800 credits finding out that none of the three problems were the same problem.
That is the expensive part. Not the distortion itself, but treating four different failures as one and paying per second to guess which one you have. This guide separates them, gives each its own fix, and puts the credit cost next to each decision.
Key Takeaways
- Distortion is four modes, not one. Identity drift, environment warp, contact breaks, and physics failures have different causes and different fixes. Applying the wrong fix costs a full re-render.
- A better-looking model does not fix physics. Across six models tested on the Physics-IQ benchmark, physical understanding was “severely limited, and unrelated to visual realism” (arxiv.org, WACV 2026).
- Prompting is the weakest lever. It helps identity and framing. It does nothing for collisions, breakage, and state changes, because the model is not failing to understand your words.
- Fix the still, then animate it. An image edit is 5 credits. A Veo 3 clip with audio at 8 seconds is 6,400. That is 1,280 image iterations for the price of one video take.
- Shorten the take before you rewrite the prompt. Error accumulates along a sequence, so the tail of a clip degrades before the head does.
- Physics failures are solved by changing the action. If the shot needs something to shatter, pour, or collide, cut around the event instead of asking a model to render it.
Why AI video distorts
Video models resolve each frame as its own problem. There is no running memory of what frame 11 decided when frame 12 gets generated, so small differences do not cancel out across a sequence. They accumulate and read as motion.
The research language for this is straightforward: the stochastic nature of sampling produces independent hallucinations in consecutive frames, which surfaces as texture flicker, jitter, and inconsistent motion (arxiv.org, 2026). A still that looks perfect paused can fail the moment it plays, because the failure is the difference between frames rather than anything inside one of them.
Two things follow from that, and they set up everything below. Anything the model has to invent will drift, because it re-invents it every frame. Anything you supply as an input has something to hold onto.
The four distortion modes, and how to tell them apart
Watch the clip once at quarter speed before you decide anything. The tells separate cleanly.
| Mode | What you see | Root cause | The fix |
|---|---|---|---|
| Identity drift | A face, garment, or product that is subtly different at the end of the clip than at the start | The subject was described in text, so the model re-interprets the description each frame | Anchor to a reference image, not an adjective |
| Environment warp | Backgrounds that bend, shelves that duplicate, straight lines that bow behind a moving subject | Background is low priority in the attention budget and gets re-solved loosely | Lock the plate before you animate it |
| Contact and anatomy | Fingers passing through a product, a grip that stops tracking the object, a jaw that changes width | High structural detail in few pixels, plus occlusion at the contact point | Repair the source still, then animate the approved frame |
| Physics failure | Objects passing through each other, things that should break staying whole, liquid that behaves wrongly | The model has no physical world model, only pixel prediction | Change the action, not the prompt |
The first three respond to better inputs. The fourth does not, and that is the part most guides get wrong.
Why a better-looking model does not fix physics
This is the finding worth internalising before you spend anything. Physics-IQ, a benchmark from INSAIT and Google DeepMind, tested whether generative video models learn physical principles from watching video. The dataset covers fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.
The result, verbatim from the paper: across Sora, Runway, Pika, Lumiere, Stable Video Diffusion, and VideoPoet, “physical understanding is severely limited, and unrelated to visual realism.” The authors close with the line that should change how you pick a model: “visual realism does not imply physical understanding” (Motamed et al., arXiv:2501.09038, WACV 2026).
Read that second clause carefully. Physical correctness and visual polish are not on the same axis. A model that renders skin, fabric, and light beautifully can still put a hand through a bottle, and paying more per second buys you the polish, not the physics.
The failure has a characteristic shape. Models rationalise events they cannot render. Ask for a glass falling off a table and current models will commonly return a glass that lands intact, because a whole glass is a more predictable set of pixels than a shattering one. The model is doing exactly what it was built to do, which is predicting likely frames, and an intact glass is the likelier set.
So the standard advice stack, write a better prompt and upgrade to the premium model, has no effect on this mode. Both levers act on the parts of the problem that were already working.
The fix that matches each mode
Identity drift: anchor, do not describe
Text is ambiguous and the model resolves the ambiguity fresh each frame. A reference image removes the ambiguity permanently.
Every video model in the DesignerBox catalogue accepts an image input. Veo 3.1 takes reference images of a scene, a character, or an object, plus first and last frame control (deepmind.google, July 2026). Sora 2 Pro accepts an input image that acts as the first frame (developers.openai.com, July 2026). Seedance 2.0 takes multi-reference input, up to 9 images plus 3 videos plus 3 audio tracks (replicate.com, July 2026).
That input changes the job from invention to preservation. For faces and recurring characters across separate clips, the deeper version of this problem has its own guide on keeping a character consistent between clips.
Environment warp: lock the plate
The background drifts because the model spends its attention on the subject. The answer is to stop asking it to generate the set at all: build the location as a still first, approve it, then use it as the start frame. Choosing a camera move the start frame can actually support is the other half of this fix, covered in how to animate a photo with AI.
This holds a set across an entire campaign rather than one clip, which is the harder and more useful version. The method is covered in locking a location across AI shots.
Contact and anatomy: repair the frame first
Hands at rest render cleanly on current models. Hands gripping, rotating, or handing over an object are where this still breaks, and that is exactly the shot product video needs.
The move is an order of operations, not a feature. Generate or shoot the still, fix it at 5 credits per edit, approve the anatomy while it is cheap, then animate the frame you approved. The tells to check and the full diagnostic are in why AI video breaks hands and faces.
Physics: change the action
There is no input that fixes this one, so the fix is directorial. Take the physical event out of the shot.
- Cut around state changes. Do not ask for the glass to break. Show the intact glass, cut, show the aftermath. Two clips that work beat one that does not.
- Show the result, not the transition. Pouring, melting, foaming, and folding are state transitions. A poured glass renders reliably. The pour does not.
- Avoid collisions entirely. Two objects meeting is the least reliable thing you can ask for. Block the shot so contact happens off-frame.
- Keep one moving element. Every additional independently moving object multiplies the physical relationships the model has to hold at once.
- Shorten the take. Error accumulates along the sequence. A clean 4 seconds beats a 10-second take that decays in the last three.
None of this is a workaround for a temporary limitation. It is shot design for the tool you have, the same way a director blocks around what a camera cannot do.
What distortion costs when you re-roll blind
Video bills per second of output, which makes guessing the most expensive habit in the workflow. Credit costs on DesignerBox, from the plan configuration:
| Operation | Credits |
|---|---|
| Generate or edit an image | 5 |
| Seedance Pro Fast, 720p, 5s | 150 |
| Kling Standard, 720p, 5s | 225 |
| Sora 2, 720p, 8s | 1,600 |
| Veo 3 with audio, 8s | 6,400 |
Three blind re-rolls of that last one is 19,200 credits. The Premium plan allocates 2,500 credits a month, so a single afternoon of guessing costs more than seven months of that tier.
Set against it: 6,400 credits buys 1,280 image edits. Diagnosing the mode on a still, where iteration is 5 credits, decides whether the workflow survives contact with a monthly budget.
Draft cheap and finish expensive. Seedance Pro Fast at 150 credits tells you whether the blocking holds. Once it does, spend the 6,400 once. If you are drafting on a vendor’s free tier instead, what free AI video plans actually issue, watermark and licence included covers the terms that decide whether the draft can ever ship.
The order to work in
Run these in sequence. Each step removes a mode before the next one gets expensive. Step zero is picking the input mode, which decides how much of the shot the model has to invent in the first place.
- Build and approve the still. Product, subject, and set, correct at full resolution, before any motion exists. This is where 5-credit iterations live.
- Check the still for anatomy and contact. Fix fingers, grip, and product geometry here. Distortion caught in a still is never a re-render.
- Design the action around physics. Remove collisions, breakage, and state transitions from the shot list. Decide where you cut.
- Draft the motion on a cheap model. 150 credits confirms whether the blocking and the camera move hold.
- Watch at quarter speed and name the mode. Use the table above. Never re-roll before you can say which of the four you are looking at.
- Finish on the expensive model, once. For prompt-side control at this stage, the seven layers of a realistic video prompt covers what to specify and what to leave alone. Camera moves have their own physics constraints, and what each model documents about executing one explains why a dolly with no depth cues comes back as a zoom.
Distortion is the expensive half of this problem. The cheaper half is everything else that makes a clip look weak, and the same ordering applies: the six AI fixes for UGC video quality are priced by when you apply them, not by which one you pick.
Switching models mid-diagnosis is the common mistake. If the mode is physics, a different model returns the same failure at a different price. DesignerBox includes 13 image and video models on one subscription, which makes swapping cheap, and that is precisely why it is worth knowing when swapping is pointless. Model specs and per-second costs sit on the model catalogue.
The wider catalogue of what breaks and what is recoverable sits in AI creative failure modes.
FAQ
Why does my AI video look fine at the start and warped at the end?
Error accumulates along the sequence. Each frame is generated from an increasingly drifted state, so the tail of a take degrades before the head does. Shorten the clip, anchor the subject with a reference image, and compare the first and last frame directly before approving. If the drift appears across separate clips rather than inside one, that is identity drift and needs a reference anchor instead.
Can a better prompt fix AI video distortion?
It fixes some of it. Prompt precision helps with framing, camera behaviour, and how a subject is described, which addresses identity drift and part of environment warp. It does nothing for physics failures. The model is not misunderstanding the words when it renders a glass landing intact, it is predicting the most likely frames. Adding “the glass shatters realistically” does not change what it can predict. The same limit applies to how people move, which is why realistic AI human movement needs structural inputs rather than better wording.
Does a more expensive model produce fewer distortions?
Not for physics. The Physics-IQ benchmark found physical understanding was unrelated to visual realism across every model tested (arxiv.org, WACV 2026). Premium models buy better texture, lighting, and motion smoothness, which does reduce the visible severity of the other three modes. Objects passing through each other happens at every price point.
How long should an AI video clip be to avoid distortion?
Shorter than the maximum, whatever the model allows. Veo 3.1 generates 8-second clips and extends from the last second (deepmind.google, July 2026). Sora 2 Pro supports 16 and 20-second generations and extends up to a 120-second total (developers.openai.com, July 2026). Seedance 2.0 runs 4 to 15 seconds, Kling 2.6 Pro 5 or 10 (replicate.com, July 2026). Generating near the cap gives error more room to accumulate. Chain shorter takes instead.
Why do objects pass through each other in AI video?
Depth ordering is resolved per frame and can flip between them. The model has no persistent representation of which object is in front, so occlusion is re-decided continuously. Physics-IQ documents this directly, listing spontaneous appearance and disappearance of objects and physically implausible interactions among the failure modes it measured (arxiv.org, WACV 2026). Block the shot so objects do not overlap, or generate them in separate passes.
Can I fix distortion after the video is generated?
Not inside the clip. Video edits are not localised the way image edits are, so a broken region means a re-render at full cost. That asymmetry is the entire argument for diagnosing on the still, where a fix is 5 credits and a mistake costs nothing.
Does the same distortion mode affect image generation?
Identity drift and contact errors do, and both are cheaper to catch there. Environment warp and physics failures are largely video-only, because both are failures of consistency between frames rather than inside one. A still has no between.
Sources
- Physical understanding “severely limited, and unrelated to visual realism” across six video models on the Physics-IQ benchmark: Motamed et al., arXiv:2501.09038 (arxiv.org, WACV 2026)
- Independent hallucinations in consecutive frames as the cause of texture flicker, jitter and inconsistent motion: (arxiv.org, 2026)
- Veo 3.1 reference images for scene, character and object, first and last frame control, 8-second clips and extension from the last second: (deepmind.google, July 2026)
- Sora 2 Pro accepting an input image as the first frame, 16 and 20-second generations, and extension to a 120-second total: (developers.openai.com, July 2026)
- Seedance 2.0 multi-reference input of 9 images, 3 videos and 3 audio tracks, plus Seedance and Kling clip durations: (replicate.com, July 2026)
- DesignerBox credit costs, plan allocations and feature gating verified against live product configuration, July 2026
Turn one product photo into the full campaign, on brand, in one workspace. See what the Ad Studio does.
Model capabilities and duration limits verified against Google DeepMind, OpenAI, and Replicate documentation as of July 2026. Benchmark findings from Physics-IQ (Motamed et al., arXiv:2501.09038, WACV 2026). Credit costs from DesignerBox plan configuration. Individual results vary.