Realistic AI video is decided by the input mode you pick before the model runs, not by the prompt you write after. Text-to-video, image-to-video, and video-to-video each protect something different: the idea, your actual product, or a real performance. Each one trades away motion freedom for fidelity. Pick the mode by what the shot cannot afford to get wrong.
You have a product photo, a deadline, and a model that will happily generate something adjacent to your product. The prompt gets rewritten four times. The clip still reads wrong, and the credits are gone.
Almost every guide on realistic AI video sends you to the prompt. Prompt craft matters, and there is a whole layer of it worth learning. It is also the last decision you make, and the smallest one. This guide covers the decision before it: what you hand the model, what each option buys, and the part no vendor guide states, which is what each option takes away.
Key Takeaways
- The input mode outranks the prompt. Text, image, and video inputs constrain the model differently. No prompt recovers a mode that was wrong for the shot.
- Starting from a photo is a trade, not an upgrade. It buys fidelity to your real product. Measured against the same architecture, image conditioning cut motion by 18.6% (arxiv.org, July 2026).
- The three modes protect different things. Text protects the idea, image protects your product, video protects a performance. Decide what the shot cannot get wrong, then pick.
- Model support varies more than the marketing suggests. Seedance 2.0 accepts up to 9 images and 3 video clips. Runway Gen-4.5 accepts text and a first frame (docs.dev.runwayml.com, July 2026).
- Iteration cost is the hidden variable. At 5 to 10 attempts per text-to-video shot, an 8-second Veo 3 render with audio runs 32,000 to 64,000 credits before you land it.
- Lock the frame at 5 credits, not the render at 6,400. Settling composition as a still is the cheapest realism lever available.
What decides whether an AI video looks real
Realism in AI video comes from constraint. A video model generates each frame by predicting what is probable, and the probable version of anything is its average. Every constraint you supply, a first frame, a reference image, a driving performance, removes a decision the model would otherwise average. The prompt is one narrow channel for constraint. The input is a much wider one.
That is why the same model returns a plastic clip and a usable one on the same day. The difference is usually not the wording. It is how much of the shot already existed before the render started.
The three input modes, and what each one protects
Every video model in the DesignerBox catalogue accepts at least one of three input modes. They are not quality tiers. They protect different things, and choosing by what needs protecting is faster than choosing by reputation.
Text-to-video protects the idea. Nothing exists yet, so nothing constrains the model, and it will invent the product, the room, and the light. Good for concept exploration, mood tests, and abstract motion where no real object has to survive. Bad for anything a customer will recognise.
Image-to-video protects the subject. Your photo becomes the first frame, so the composition, the light, and the object are settled before motion begins. This is the mode that makes the output your product rather than a lookalike. It is the default for commercial work for exactly that reason.
Video-to-video protects the performance. A driving clip supplies the motion, and the model changes the style, environment, or subject around it. Good for human action that generators still struggle to invent, since real motion cannot be averaged away if it was supplied.
The pattern underneath: each mode hands the model more of the answer, and takes away more of its freedom. That freedom is not always something you want back.
What the research actually says about starting from a photo
Here the guides and the measurements disagree, and it is worth knowing which one you are trusting.
The standard claim is that image-to-video is simply more consistent than text-to-video. Runway’s own guide states that image input gives “much more consistent results than text-only generation because the composition, lighting, and subject are already locked in” (runway.com, published November 2025). No provider publishes a measurement behind that comparison, and the published research is split.
What supports it. Holding the base model, training data, and parameters constant, Meta’s Emu Video found human evaluators preferred image-factorized generation over direct text-to-video 70.5% on quality and 63.3% on faithfulness (arxiv.org, accessed July 2026). One caveat matters: the conditioning image there was generated by the model from the text, not uploaded by a user. It supports routing generation through an image. It does not measure your product photo.
What cuts against it. Running Wan 2.1 in both modes on the same architecture, researchers fed the text model’s own first frame into the image model and measured the result:
| Metric | Text-to-video | Image-to-video | Change |
|---|---|---|---|
| Dynamic degree (motion) | 39.4 | 32.1 | -18.6% |
| Subject consistency | 97.1 | 95.2 | -2.0% |
| Aesthetic quality | 59.5 | 56.8 | -4.6% |
The paper’s own summary is that image-to-video models “frequently produce much more static videos compared to their text-to-video counterparts, often adhering too closely to the reference image” (arxiv.org, accessed July 2026). The VBench++ benchmark scores image-to-video on separate axes for the same reason, noting that a model’s ability to align closely with the input image “may limit its capacity to maintain temporal consistency across video frames” (arxiv.org, accessed July 2026). ByteDance’s own Seedance technical report says adding an image condition “introduces challenges in preserving character and background” (arxiv.org, accessed July 2026).
What this means for a real brief. Starting from your product photo is still the right call for commercial work, and the reason is fidelity to your actual product, not smoother video. You are buying the guarantee that the bottle on screen is your bottle. You pay for it in motion range, which is why image-driven clips often feel static and why the fix is a shorter take with less asked of it, rather than a longer prompt. Matching the motion request to what the still can actually support is its own skill, covered in how to animate a photo with AI. Anyone selling image-to-video as a free upgrade has not measured it.
Which models accept which inputs
Support varies more than the category’s marketing suggests. Verified against provider documentation, July 2026:
| Model | Text | First frame | Last frame | Reference images | Video input |
|---|---|---|---|---|---|
| Veo 3.1 | Yes | Yes | Yes | Up to 3 | Not documented |
| Veo 3.1 Fast | Yes | Yes | Yes | Up to 3 | Not documented |
| Sora 2 Pro | Yes | Yes | Not documented | Reusable subject asset | Not documented |
| Seedance 2.0 | Yes | Yes | Yes | Up to 9 | Up to 3 clips |
| Kling 2.6 Pro | Yes | Yes | Yes | Not supported | Motion Control |
| Runway Gen-4.5 | Yes | Yes | No | No | No |
Three things worth reading off that table.
Seedance 2.0 accepts the widest input set, taking up to 9 images and 3 video clips in one instruction (seed.bytedance.com, accessed July 2026). If a shot needs several references held at once, that is the model built for it.
Kling 2.6 Pro has no reference-image support, but its Motion Control feature drives a character image from an uploaded action video (kling.ai, accessed July 2026). That is a video-to-video lever reached from a different direction.
Runway Gen-4.5 covers text and a first frame, with durations of 2 to 10 seconds (docs.dev.runwayml.com, accessed July 2026). Runway handles video editing and performance transfer with separate models rather than folding them into Gen-4.5, so a first frame is the constraint available on that one.
One freshness note that applies to every guide in this category, this one included. Runway’s tips page is dated November 2025 and Gen-4.5 shipped that December, gaining first-frame image input in January 2026 (runway.com/changelog, accessed July 2026). Input-mode support moves on a monthly cycle. Check the provider’s current documentation before you plan a shot around a capability, including the table above.
How to pick the input mode for your shot
Work the question in this order. The first yes is your answer.
- Does a real product have to be recognisable on screen? Use image-to-video, starting from the actual product photo. Nothing else guarantees the object is yours.
- Does a specific human action have to read correctly? Use video-to-video or a motion-control feature, and supply the performance. Complex human motion is the thing generators average hardest.
- Does a character or location have to hold across several clips? Use reference images, which means Veo 3.1 or Seedance 2.0. Details in keeping characters consistent across AI video clips and holding locations across shots.
- Is nothing real yet? Text-to-video, and treat the output as a concept frame rather than a deliverable. Cheaper still, settle the framing as a storyboard first, covered in script to storyboard.
For most brand work the answer is question one, and the honest version of that answer is that you already own the input. The product photo on your PDP is the constraint. DesignerBox’s image to video tool runs that path directly, and the full model catalogue shows which of the 13 models accepts what.
What each mode costs before it lands
This is the part vendor guides leave out, and it changes which mode a budget can afford.
Runway’s guide puts text-to-video at 5 to 10 attempts to get close to what you want (runway.com, November 2025). Price that against DesignerBox’s own video rates, where video is billed as credits per second of output:
| Example clip | Credits | Per second |
|---|---|---|
| Seedance Pro Fast, 720p, 5s | 150 | 30 |
| Kling Standard, 720p, 5s | 225 | 45 |
| Sora 2, 720p, 8s | 1,600 | 200 |
| Veo 3 + Audio, 8s | 6,400 | 800 |
At 5 to 10 attempts, an 8-second Veo 3 render with audio costs 32,000 to 64,000 credits to land. The Ultra plan includes 8,000 credits a month. That single shot is four to eight months of allocation, spent guessing.
The same shot approached from a still costs 5 credits per image. Twenty attempts at the composition is 100 credits, roughly one and a half percent of one Veo render. You settle framing, light, and the product’s appearance while every attempt is cheap, then spend the video budget once on a take you have already seen.
That produces the rule the whole category runs on: draft cheap, finish expensive. Block the shot on Seedance Pro Fast, lock what you can as a still, and buy the expensive render last. Full per-second figures are in AI video generation cost, and the plan-by-plan version is in how many clips your credits buy.
Worth noting on plans: AI video generation requires the Premium tier or higher, and the commercial licence starts at Pro.
What the input cannot fix
Choosing the right mode removes a class of problems. It does not remove all of them, and knowing the line saves re-rolls.
A correct input still leaves you exposed to frame-to-frame instability, since models generate each frame with no memory of the last. Identity drift, environment warp, and contact and anatomy breaks are separate failures with separate fixes, worked through in how to avoid distortions in AI video and why hands and faces break.
A correct input also does not supply weight. A model can hold your product perfectly and still show someone lifting it as though it were empty, because effort gets averaged out of predicted motion. That is its own problem, covered in realistic AI human movement.
And once the mode is right, the prompt does start to matter. It is the last layer, not the first, and the seven things worth stating are in how to write realistic AI video prompts. The order those things go in changes per model, which is why each provider documents a different prompt structure. DesignerBox’s failure modes guide maps what the category still gets wrong in 2026.
One habit pays for itself here: write down which input mode, which model, and which settings produced a take that worked. A production diary turns a lucky render into a repeatable one, which matters more as the shot count grows.
Start from the shot you already own
The input decision is made before any prompt exists, and for brand work it is usually already made for you. You have a product photo. That photo is the strongest constraint available, and using it is what keeps the output your product instead of something adjacent to it.
DesignerBox includes all 13 models on one subscription, so the input mode can be chosen per shot without a second bill or a second login. One photo goes in and the campaign comes out of the same workspace.
FAQ
What is the best input mode for realistic AI video?
For commercial work, image-to-video starting from your real product photo. It fixes the composition, light, and object before motion begins, which is what makes the output your product rather than a lookalike. Text-to-video suits concept exploration where nothing real has to survive. Video-to-video suits complex human action, because a supplied performance cannot be averaged away.
Is image-to-video really more consistent than text-to-video?
Not straightforwardly. No provider publishes a measurement behind that comparison, and the research is split. Meta’s Emu Video found evaluators preferred image-factorized generation 70.5% on quality, while a same-architecture test of Wan 2.1 found image conditioning cut motion by 18.6% and subject consistency by 2.0% (arxiv.org, July 2026). The reliable benefit is fidelity to your actual input, not smoother video.
Why does my AI video look static when I start from a photo?
Because image conditioning constrains the model toward the reference frame. Benchmarked on the same architecture, image-to-video reduced dynamic degree from 39.4 to 32.1, and the researchers described these models as “adhering too closely to the reference image” (arxiv.org, July 2026). Ask for less motion in a shorter take rather than fighting it, or supply the motion directly with a video input.
Which AI video models accept reference images?
Veo 3.1 and Veo 3.1 Fast accept up to 3 reference images (ai.google.dev, July 2026). Seedance 2.0 accepts up to 9 images plus 3 video clips in one instruction (seed.bytedance.com, July 2026). Kling 2.6 Pro does not support reference images but offers Motion Control from an uploaded action video (kling.ai, July 2026). Runway Gen-4.5 accepts text and a first frame.
How many attempts does an AI video shot take?
Runway puts text-to-video at 5 to 10 attempts (runway.com, November 2025). Starting from a settled still cuts that sharply, because composition and lighting stop being variables. The cost difference is the point: an image is 5 credits, while an 8-second Veo 3 render with audio is 6,400, so guessing at the still stage is roughly a thousand times cheaper than guessing at the render stage.
Do I have to disclose AI-generated video in ads?
It depends on the platform. TikTok requires all ads with AI-generated content to carry the AIGC label or an equivalent disclosure, and states that ads without it may be rejected or restricted, while exempting minor edits such as background removal and colour adjustment (ads.tiktok.com, updated April 2026). Meta applies labels automatically via detection for commercial ads, with active disclosure required for social issue, election, and political ads (facebook.com/business, July 2026). Google Ads requires disclosure for election ads and makes the label optional for commercial ones (support.google.com, July 2026). Verify the current policy before a campaign ships.
Does the input mode fix distorted hands or faces?
No. A correct input settles what is in the frame, not how stably the model redraws it. Hands, faces, and object contact break because each frame is generated without memory of the previous one. Starting from a still where the anatomy is already correct removes much of the risk, but it is a separate fix from the input-mode decision.
Sources
All accessed July 2026.
- Veo 3.1 and Veo 3.1 Fast input modes, first and last frame support, and up to three reference images: (docs.cloud.google.com and ai.google.dev, July 2026)
- Sora 2 Pro input reference acting as the first frame, and reusable subject assets: (developers.openai.com, July 2026)
- Seedance 2.0 accepting up to 9 images and 3 video clips in one instruction: (seed.bytedance.com, July 2026)
- Kling 2.6 capability matrix and Motion Control driven by an uploaded action video: (kling.ai, July 2026)
- Runway Gen-4.5 text and image-to-video support, 2 to 10 second durations, and the January 2026 first-frame changelog entry: (docs.dev.runwayml.com and runway.com/changelog, July 2026)
- Attempt counts and the image-to-video consistency claim: (runway.com/resources/tips-make-realistic-ai-videos, published November 2025)
- Emu Video human preference rates of 70.5% on quality and 63.3% on faithfulness: (arxiv.org/abs/2311.10709, July 2026)
- Wan 2.1 same-architecture comparison showing dynamic degree, subject consistency and aesthetic quality changes: (arxiv.org/abs/2506.08456, July 2026)
- VBench++ on image alignment limiting temporal consistency: (arxiv.org/abs/2411.13503, July 2026)
- Seedance technical report on image conditioning and character preservation: (arxiv.org/abs/2506.09113, July 2026)
- TikTok, Meta and Google Ads AI disclosure requirements for advertisers: (ads.tiktok.com, facebook.com/business and support.google.com, July 2026)
- DesignerBox credit costs, plan allocations, model catalogue and feature gating verified against live product configuration, July 2026
Model input-mode support verified from provider documentation at docs.cloud.google.com, ai.google.dev, developers.openai.com, seed.bytedance.com, kling.ai and docs.dev.runwayml.com as of July 2026. Research figures from arXiv preprints as cited. DesignerBox pricing and credit costs from the product’s own plans. Individual results vary.