Avoid distortions in AI video by naming which of four failures you have before you re-roll. Identity drift needs a reference image. Environment warp needs the background locked as an approved still. Contact and anatomy breaks are repaired in the still, then animated. Physics failures need the action changed, because no prompt and no premium model fixes them.
You have a six-second clip of a model picking up your bottle. The face is right at second one and someone else’s by second five. The shelf behind her has grown a second shelf. The bottle passes through her thumb. You re-roll it twice on the premium model, three takes in total, and discover that none of the three problems were the same problem.
Small errors between frames add up across a clip, so three unrelated failures can share one take. That is the expensive part: treating four failures as one and paying for a full clip to guess which one you have. This guide separates them, gives each its own fix, and puts the cost next to each decision.
Key Takeaways
- Distortion comes in four modes. Identity drift, environment warp, contact breaks, and physics failures have different causes and different fixes. Applying the wrong fix costs a full re-render.
- A better-looking model does not fix physics. Across six models tested on the Physics-IQ benchmark, physical understanding was “severely limited, and unrelated to visual realism” (arxiv.org, WACV 2026).
- Prompting is the weakest lever. It helps identity and framing. It does nothing for collisions, breakage, and state changes, because those failures have nothing to do with your words.
- Fix the still, then animate it. An image edit costs less than even the cheapest 8-second clip, and an 8-second clip costs 40 to 560 credits, depending on the model.
- Shorten the take before you rewrite the prompt. Error accumulates along a sequence, so the tail of a clip degrades before the head does.
- Physics failures are solved by changing the action. If the shot needs something to shatter, pour, or collide, cut around the event instead of asking a model to render it.
Why AI video distorts, and how to avoid it
Video models have to keep every frame in agreement with the frames around it, and they do not always manage it. Small differences between frames do not cancel out across a sequence. They accumulate and read as motion.
Research on the extreme case is straightforward. When image diffusion models are applied to video one frame at a time, the randomness of sampling produces independent hallucinations in consecutive frames, which surfaces as texture flicker, jitter, and inconsistent motion (arxiv.org, October 2025). Video models are built to reduce that, but the same kind of error still appears. A still that looks perfect paused can fail the moment it plays, because the failure is the difference between frames rather than anything inside one of them.
Two things follow from that, and they set up everything below. Anything the model has to invent will drift, because it has to hold that invention steady in every frame. Anything you supply as an input has something to hold onto. So the first way to avoid distortion is to supply more and ask the model to invent less.
The four distortion modes, and how to tell them apart
To avoid distortions in AI video, name which of four modes you have, then apply the fix for that mode. Watch the clip once at quarter speed before you decide anything. The tells separate cleanly.
| Mode | What you see | Root cause | The fix |
|---|---|---|---|
| Identity drift | A face, garment, or product that is subtly different at the end of the clip than at the start | The subject was described in text, so the model re-interprets the description each frame | Anchor to a reference image instead of an adjective |
| Environment warp | Backgrounds that bend, shelves that duplicate, straight lines that bow behind a moving subject | Background is low priority in the attention budget and gets re-solved loosely | Lock the plate before you animate it |
| Contact and anatomy | Fingers passing through a product, a grip that stops tracking the object, a jaw that changes width | High structural detail in few pixels, plus occlusion at the contact point | Repair the source still, then animate the approved frame |
| Physics failure | Objects passing through each other, things that should break staying whole, liquid that behaves wrongly | The model predicts pixels and has no physical world model | Change the action instead of the prompt |
The first three respond to better inputs. The fourth does not, so it needs a different kind of fix.
Why a better-looking model does not fix physics
A better-looking or more expensive video model does not fix physics failures. Physics-IQ, a benchmark from INSAIT and Google DeepMind, tested whether generative video models learn physical principles from watching video. The dataset covers fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.
The result, verbatim from the paper: across Sora, Runway, Pika, Lumiere, Stable Video Diffusion, and VideoPoet, “physical understanding is severely limited, and unrelated to visual realism.” The authors close with the line that should change how you pick a model: “visual realism does not imply physical understanding” (Motamed et al., arXiv:2501.09038, WACV 2026).
Read that second clause carefully. Physical correctness and visual polish sit on different axes. A model that renders skin, fabric, and light beautifully can still put a hand through a bottle, and paying more per second buys polish while the physics stays where it was. The six models in the paper are older than today’s releases, so treat the finding as a pattern to test on your own shots.
The failure has a characteristic shape. Models rationalize events they cannot render. Ask for a glass falling off a table and current models will commonly return a glass that lands intact, because a whole glass is a more predictable set of pixels than a shattering one. The model is doing exactly what it was built to do, which is predicting likely frames, and an intact glass is the likelier set.
So the two usual levers, a better prompt and a premium model, have no effect on this mode. Both levers act on the parts of the problem that were already working.
The fix that matches each mode
Identity drift: anchor, do not describe
Text is ambiguous and the model resolves the ambiguity fresh each frame. A reference image removes the ambiguity permanently.
Most current video models take an image input. Veo 3.1 and Veo 3.1 Fast take up to three reference images of a single person, character or product, and all three Veo 3.1 models take a first and last frame (ai.google.dev, September 2026). Sora 2 Pro accepted an input image that acted as the first frame (developers.openai.com, September 2026). OpenAI removed Sora 2 and Sora 2 Pro from its API on 24 September 2026 (OpenAI API deprecations, accessed September 2026). Seedance 2.0 takes up to 9 images, 3 video clips and 3 audio clips as reference (seed.bytedance.com, September 2026).
That input changes the job from invention to preservation. For faces and recurring characters across separate clips, the deeper version of this problem has its own guide on keeping a character consistent between clips.
Environment warp: lock the plate
The background drifts because the model spends its attention on the subject. The answer is to stop asking it to generate the set at all: build the location as a still first, approve it, then use it as the start frame. Choosing a camera move the start frame can support is the other half of this fix, covered in how to animate a photo with AI.
This holds a set across an entire campaign rather than one clip, which is the harder and more useful version. The method is covered in locking a location across AI shots.
Contact and anatomy: repair the frame first
Hands at rest render cleanly on current models. Hands gripping, rotating, or handing over an object are where this still breaks, and that is exactly the shot product video needs.
The move is an order of operations. Generate or shoot the still, fix it with an image edit, approve the anatomy while it is cheap, then animate the frame you approved. The tells to check and the full diagnostic are in why AI video breaks hands and faces.
Physics: change the action
There is no input that fixes this one, so the fix is directorial. Take the physical event out of the shot.
- Cut around state changes. Do not ask for the glass to break. Show the intact glass, cut, show the aftermath. Two clips that work beat one that does not, and the multi-track video timeline is where they get joined.
- Show the result instead of the transition. Pouring, melting, foaming, and folding are state transitions. A poured glass renders reliably. The pour does not.
- Avoid collisions entirely. Two objects meeting is the least reliable thing you can ask for. Block the shot so contact happens off-frame.
- Keep one moving element. Every additional independently moving object multiplies the physical relationships the model has to hold at once.
- Shorten the take. Error accumulates along the sequence. A clean 4 seconds beats a 10-second take that decays in the last three.
This is shot design for the tool you have, the same way a director blocks around what a camera cannot do.
The cost of a blind re-roll
Every re-roll is a full video run. In DesignerBox, an 8-second clip costs 40 to 560 credits, depending on the model. The cost is shown before the run, which makes blind re-rolling a decision rather than an accident.
An image edit costs less than even the cheapest 8-second clip. So diagnosing the mode on a still is what keeps the work inside a monthly plan. There is a free plan, and it cannot make video. Uploading your own photos and the commercial license start on the Pro plan. AI video, virtual try-on, upscaling, the image editor and the video editor start on the Premium plan. Plans and credits are on the pricing page. The budget question is which model each take deserves. The same reject math applies to stills, and the AI image retry rate puts a number on it.
Draft cheap and finish expensive. A lower-cost video model tells you whether the blocking holds. Once it does, run the premium model once. Fewer re-rolls also shorten the schedule, which how to make AI videos fast covers. If you are drafting on a vendor’s free tier instead, what free AI video plans issue, watermark and license included covers the terms that decide whether the draft can ever ship.
The order to work in
Run these in sequence. Each step removes a mode before the next one gets expensive. Step zero is picking the input mode, which decides how much of the shot the model has to invent in the first place.
- Build and approve the still. Product, subject, and set, correct at full resolution, before any motion exists. This is where the cheap image edits happen. Clear it against the four gates that decide what ships before you animate.
- Check the still for anatomy and contact. Fix fingers, grip, and product geometry here. Distortion caught in a still is never a re-render.
- Design the action around physics. Remove collisions, breakage, and state transitions from the shot list. Decide where you cut.
- Draft the motion on a cheap model. One run on a lower-cost model confirms whether the blocking and the camera move hold at eight seconds.
- Watch at quarter speed and name the mode. Use the table above. Never re-roll before you can say which of the four you are looking at.
- Finish on the expensive model, once. For prompt-side control at this stage, the seven layers of a realistic video prompt covers what to specify and what to leave alone. Camera moves have their own physics constraints, and what each model documents about executing one explains why a dolly with no depth cues comes back as a zoom.
Distortion is the expensive half of this problem. The cheaper half is everything else that makes a clip look weak, and the same ordering applies: the six AI fixes for UGC video quality are priced by when you apply them.
Everything this order needs sits in one place, in one subscription. That is the image editor for the still, the video editor for the cut, brand rules, Assets and the model list. A diagnosis never moves between tools.
Switching models mid-diagnosis is the common mistake. If the mode is physics, a different model returns the same failure at a different price. In DesignerBox the model is one step in a workflow, so you change it in that step and the rest stays the same. That is exactly why it is worth knowing when a change is pointless. The model list shows what each model does.
The other half of a consistent set is the brand itself. You set the brand once, and the workflow reads it on every run, so the light, the palette and the framing do not drift while you are chasing a physics failure. That record lives in brand rules.
Build the shot once as a workflow, approve it, then run it on the next product and the one after that. Start from a template, add your brand and your products, and run it. The cost is shown before the run. See the templates.
FAQ
Why does my AI video look fine at the start and warped at the end?
Error accumulates along the sequence. Each frame is generated from an increasingly drifted state, so the tail of a take degrades before the head does. Shorten the clip, anchor the subject with a reference image, and compare the first and last frame directly before approving. If the drift appears across separate clips rather than inside one, that is identity drift and needs a reference anchor instead.
Can a better prompt fix AI video distortion?
It fixes some of it. Prompt precision helps with framing, camera behavior, and how a subject is described, which addresses identity drift and part of environment warp. It does nothing for physics failures. When the model renders a glass landing intact, it is predicting the most likely frames. Adding “the glass shatters realistically” does not change what it can predict. The same limit applies to how people move, which is why realistic AI human movement needs structural inputs rather than better wording.
Does a more expensive model produce fewer distortions?
Not for physics. The Physics-IQ benchmark found physical understanding was unrelated to visual realism across every model tested (arxiv.org, WACV 2026). Premium models buy better texture, lighting, and motion smoothness, which does reduce the visible severity of the other three modes. Objects passing through each other happens at every price point.
How long should an AI video clip be to avoid distortion?
Shorter than the maximum, whatever the model allows. Veo 3.1 generates up to 8 seconds, and in the Gemini API it extends a clip 7 seconds at a time (ai.google.dev, September 2026). Until its API removal on 24 September 2026, Sora 2 Pro generated up to 20 seconds and extended to a 120-second total (developers.openai.com, September 2026). Seedance 2.0 runs 4 to 15 seconds, and Kling 2.6 makes 5-second or 10-second clips (docs.byteplus.com and kling.ai, September 2026). Generating near the cap gives error more room to accumulate. Chain shorter takes instead.
Why do objects pass through each other in AI video?
Depth ordering is resolved per frame and can flip between them. The model has no persistent representation of which object is in front, so occlusion is re-decided continuously. Physics-IQ documents this directly, listing spontaneous appearance and disappearance of objects and physically implausible interactions among the failure modes it measured (arxiv.org, WACV 2026). Block the shot so objects do not overlap, or generate them in separate passes.
Can I fix distortion after the video is generated?
Not inside the clip. Video edits are not localized the way image edits are, so a broken region means a re-render at full cost. That asymmetry is the entire argument for diagnosing on the still, where a fix is one image edit and costs less than any 8-second clip.
Does the same distortion mode affect image generation?
Identity drift and contact errors do, and both are cheaper to catch there. Environment warp and physics failures are largely video-only, because both are failures of consistency between frames rather than inside one. A still has no between.
Sources
- Physical understanding “severely limited, and unrelated to visual realism” across six generative video models on the Physics-IQ benchmark: Motamed et al., arXiv:2501.09038 (arxiv.org, WACV 2026)
- Independent hallucinations in consecutive frames when image diffusion models are applied to video frame by frame: arXiv:2510.25420 (arxiv.org, October 2025)
- Veo 3.1 reference images, first and last frame control, 8-second clips and 7-second extension in the Gemini API: ai.google.dev, accessed October 2026
- Sora 2 Pro input image as the first frame, clip lengths and extension to a 120-second total: OpenAI video generation guide, September 2026
- Sora 2 and Sora 2 Pro removal from OpenAI’s API on 24 September 2026: OpenAI API deprecations, accessed October 2026
- Seedance 2.0 reference input of 9 images, 3 video clips and 3 audio clips: seed.bytedance.com, September 2026
- Seedance 2.0 and Kling 2.6 clip durations: docs.byteplus.com and kling.ai, September 2026
- DesignerBox plans, credits and feature gating: DesignerBox pricing page (designerbox.ai/pricing), September 2026
Model capabilities and duration limits verified against Google, OpenAI, ByteDance and Kuaishou documentation as of September 2026. Google’s Veo guide, OpenAI’s deprecations page and both arXiv papers were re-checked on 2 October 2026. Benchmark findings from Physics-IQ (Motamed et al., arXiv:2501.09038, WACV 2026). Individual results vary.