Learning how to make realistic AI videos starts with one choice. Pick the input mode before you write the prompt. Text-to-video, image-to-video and video-to-video each protect something different: the idea, your actual product, or a real performance. Each one buys fidelity by giving up motion range. On a same-architecture test, image conditioning cut measured motion by 18.6%. Pick the mode by what the shot cannot afford to get wrong.
You have a product photo, a deadline, and a model that will happily generate something adjacent to your product. The prompt gets rewritten four times. The clip still reads wrong, and the credits are gone.
Prompt craft matters, and there is a whole layer of it worth learning. It is also the last decision you make, and the smallest one. This guide covers the decision before it: what you hand the model, what each option buys, and what each option takes away.
Key Takeaways
- The input mode outranks the prompt. Text, image, and video inputs constrain the model differently. No prompt recovers a mode that was wrong for the shot.
- Starting from a photo is a trade, not an upgrade. It buys fidelity to your real product. Measured against the same architecture, image conditioning cut motion by 18.6% (arxiv.org, October 2026).
- The three modes protect different things. Text protects the idea, image protects your product, video protects a performance. Decide what the shot cannot get wrong, then pick.
- Iteration cost is the hidden variable. Runway puts text-to-video at 5 to 10 attempts per shot, and every attempt is a paid video run.
- Settle the frame as a still first. Attempts at the composition cost far less as stills than as video takes.
How to make realistic AI videos with constraints
Realistic AI videos come from constraint. A video model generates each frame by predicting what is probable, and the probable version of anything is its average. Every constraint you supply, a first frame, a reference image, a driving performance, removes a decision the model would otherwise average. The prompt is one narrow channel for constraint. The input is a much wider one.
That is why the same model returns a plastic clip and a usable one on the same day. The difference is usually how much of the shot already existed before the render started.
The three input modes, and what each one protects
Every video model in the DesignerBox catalog accepts at least one of three input modes. They are not quality tiers. They protect different things, and choosing by what needs protecting is faster than choosing by reputation.
Text-to-video protects the idea. Nothing exists yet, so nothing constrains the model, and it will invent the product, the room, and the light. Good for concept exploration, mood tests, and abstract motion where no real object has to survive. Bad for anything a customer will recognize.
Image-to-video protects the subject. Your photo becomes the first frame, so the composition, the light, and the object are settled before motion begins. This is the mode that makes the output your product rather than a lookalike. It is the default for commercial work for exactly that reason.
Video-to-video protects the performance. A driving clip supplies the motion, and the model changes the style, environment, or subject around it. Good for human action that generators still struggle to invent, since real motion cannot be averaged away if it was supplied.
The pattern underneath: each mode hands the model more of the answer, and takes away more of its freedom. That freedom is not always something you want back.
What the research says about starting from a photo
The research on starting from a photo is split. Guides call image-to-video more consistent, while a same-architecture test measured 18.6% less motion.
The standard claim is that image-to-video is more consistent than text-to-video, full stop. Runway’s own guide states that image-to-video gives “much more consistent results than text-only generation because the composition, lighting, and subject are already locked in” (runway.com, published 11 November 2025, accessed September 2026). No provider publishes a measurement behind that comparison, and the published research is split.
What supports it. Holding the base model, training data, and parameters constant, Meta’s Emu Video found human evaluators preferred image-factorized generation over direct text-to-video 70.5% on quality and 63.3% on faithfulness (arxiv.org, accessed October 2026). One caveat matters: the conditioning image there was generated by the model from the text, not uploaded by a user. It supports routing generation through an image. It does not measure your product photo.
What cuts against it. Running Wan 2.1 in both modes on the same architecture, researchers fed the text model’s own first frame into the image model and measured the result:
| Metric | Text-to-video | Image-to-video | Change |
|---|---|---|---|
| Dynamic degree (motion) | 39.4 | 32.1 | -18.6% |
| Subject consistency | 97.1 | 95.2 | -2.0% |
| Aesthetic quality | 59.5 | 56.8 | -4.6% |
The paper’s own summary is that image-to-video models “frequently produce much more static videos compared to their T2V [text-to-video] counterparts, often adhering too closely to the reference image” (arXiv:2506.08456, accessed October 2026). The VBench++ benchmark scores image-to-video on separate axes for the same reason, noting that a model’s ability to align closely with the input image “may limit its capacity to maintain temporal consistency across video frames” (arXiv:2411.13503, accessed October 2026). ByteDance’s Seedance 1.0 technical report, describing image-to-video in general, says an image condition “introduces challenges in preserving character and background” (arXiv:2506.09113, accessed October 2026).
What this means for a real brief. Starting from your product photo is still the right call for commercial work, and the reason is fidelity to your actual product, not smoother video. You are buying the certainty that the bottle on screen is your bottle. You pay for it in motion range, which is why image-driven clips often feel static and why the fix is a shorter take with less asked of it, rather than a longer prompt. Matching the motion request to what the still can support is its own skill, covered in how to animate a photo with AI. Anyone selling image-to-video as a free upgrade has not measured it.
Which models accept which inputs
Support varies more than the category’s marketing suggests. Verified against provider documentation, September 2026. Sora 2 Pro is not in the table: OpenAI removed Sora 2 and Sora 2 Pro from its API on 24 September 2026 (OpenAI API deprecations, accessed October 2026), and DesignerBox retired it from its roster on 18 September 2026. Sora alternatives by job covers what replaces it.
| Model | Text | First frame | Last frame | Reference images | Video input |
|---|---|---|---|---|---|
| Veo 3.1 | Yes | Yes | Yes | Up to 3 | Not documented |
| Veo 3.1 Fast | Yes | Yes | Yes | Up to 3 | Not documented |
| Seedance 2.0 | Yes | Yes | Yes | Up to 9 | Up to 3 clips |
| Kling 2.6 Pro | Yes | Yes | Yes, at 1080p, silent | Not supported | Motion Control |
| Runway Gen-4.5 | Yes | Yes | No | No | No |
Four things the table shows.
Among these models, Seedance 2.0 accepts the widest input set, taking up to 9 images, 3 video clips and 3 audio clips in one request (seed.bytedance.com, accessed September 2026). It also supports first and last frame input (docs.byteplus.com, September 2026). If a shot needs several references held at once, that is the model built for it. ByteDance’s newer Seedance 2.5 takes up to 50 references, and DesignerBox does not carry it.
Kling 2.6 Pro has no reference-image support, but its Motion Control feature drives a character image from an uploaded action video (kling.ai, accessed September 2026). That is a video-to-video lever reached from a different direction. Kling 3.0 does take references: up to 3 Elements per clip, each built from a video or 2 to 4 images.
Runway Gen-4.5 covers text and a first frame, with durations of 2 to 10 seconds (docs.dev.runwayml.com, accessed September 2026). Runway handles video editing and performance transfer with separate models rather than folding them into Gen-4.5, so a first frame is the constraint available on that one. Runway alternatives by job lists the options when a shot needs more than that.
One freshness note that applies to every guide in this category, this one included. Runway’s tips page is dated 11 November 2025. Gen-4.5 was announced on 1 December 2025 and gained first-frame image input on 21 January 2026 (runway.com/changelog, accessed September 2026). Google now names Gemini Omni Flash its default video model and keeps Veo 3.1 for scene extension and last-frame control (ai.google.dev, October 2026). Google lists 22 October 2026 as the earliest shutdown date for the Veo 3.1 preview models in the Gemini API, and names Gemini Omni Flash as the replacement (Gemini API deprecations, October 2026). Input-mode support moves on a monthly cycle. Check the provider’s current documentation before you plan a shot around a capability, including the table above.
How to pick the input mode for your shot
Work the question in this order. The first yes is your answer.
- Does a real product have to be recognizable on screen? Use image-to-video, starting from the actual product photo. Nothing else holds the object to the one you sell.
- Does a specific human action have to read correctly? Use video-to-video or a motion-control feature, and supply the performance. Complex human motion is the thing generators average hardest.
- Does a character or location have to hold across several clips? Use reference images, which means Veo 3.1, Seedance 2.0 or Kling 3.0. Details in keeping characters consistent across AI video clips and holding locations across shots.
- Is nothing real yet? Text-to-video, and treat the output as a concept frame rather than a deliverable. Cheaper still, settle the framing as a storyboard first, covered in script to storyboard.
For most brand work the answer is question one, and the honest version of that answer is that you already own the input. The product photo on your PDP is the constraint, and the rules for a product clip that works are in AI product video. DesignerBox’s image to video app runs that path directly. DesignerBox runs video models, and the model list names them.
Cost of attempts per shot
The number of attempts a shot needs changes which mode a budget can afford.
Runway’s guide says “You’ll likely need 5-10 attempts to get close to what you want” with text-to-video (runway.com, accessed September 2026). Every one of those attempts is a paid video run. In DesignerBox an 8-second clip costs 40 to 560 credits, depending on the model. So the model you route the guessing to matters more than the guessing does. Ten attempts on a low-cost model and ten on a premium model are very different bills for the same shot.
The same shot approached from a still costs far less per attempt. You settle framing, light and the product’s appearance while every attempt is cheap, then spend the video budget once on a take you have already seen.
That produces the rule the whole category runs on: draft cheap, finish expensive. Block the shot on a low-cost model, lock what you can as a still, and run the expensive model last. Every run shows its cost before you start it, which is what makes the rule enforceable rather than aspirational. What drives the cost of a shot is in AI video generation cost.
Worth noting on plans. Uploading your own photos and the commercial license start on the Pro plan. AI video, virtual try-on, upscaling, the image editor and the video editor start on the Premium plan. Plans and credits are on the pricing page.
What the input cannot fix
Choosing the right mode removes a class of problems. It does not remove all of them, and knowing the line saves re-rolls.
A correct input still leaves you exposed to frame-to-frame instability, since models generate each frame with no memory of the last. Identity drift, environment warp, and contact and anatomy breaks are separate failures with separate fixes, worked through in how to avoid distortions in AI video and why hands and faces break.
A correct input also does not supply weight. A model can hold your product perfectly and still show someone lifting it as though it were empty, because effort gets averaged out of predicted motion. That is its own problem, covered in realistic AI human movement.
And once the mode is right, the prompt does start to matter. It is the last layer, not the first, and the seven things worth stating are in how to write realistic AI video prompts. The order those things go in changes per model, which is why each provider documents a different prompt structure. The one thing you should not repeat in every prompt is your brand: DesignerBox holds brand rules the workflow reads on every run.
One habit pays for itself here: write down which input mode, which model, and which settings produced a take that worked. That note turns a lucky take into a repeatable one, which matters more as the shot count grows.
The shot you already own
The input decision is made before any prompt exists, and for brand work it is usually already made for you. You have a product photo. That photo is the strongest constraint available, and using it is what keeps the result your product instead of something adjacent to it.
Do that once properly and it stops being a per-shot decision. DesignerBox is AI creative production for brands and agencies. You set the input mode, the model and the settings once as a workflow, with your brand and your products. A saved workflow runs the same way on the next product and the one after that. The still, the take and the edit are all part of DesignerBox. The full workflow from the first product photo to the finished ad, in one subscription. Start from the video ad templates.
FAQ
What is the best input mode for realistic AI video?
For commercial work, image-to-video starting from your real product photo. It fixes the composition, light, and object before motion begins, which is what makes the output your product rather than a lookalike. Text-to-video suits concept exploration where nothing real has to survive. Video-to-video suits complex human action, because a supplied performance cannot be averaged away.
Is image-to-video really more consistent than text-to-video?
Not straightforwardly. No provider publishes a measurement behind that comparison, and the research is split. Meta’s Emu Video found evaluators preferred image-factorized generation 70.5% on quality, while a same-architecture test of Wan 2.1 found image conditioning cut motion by 18.6% and subject consistency by 2.0% (arxiv.org, October 2026). The reliable benefit is fidelity to your actual input, not smoother video.
Why does my AI video look static when I start from a photo?
Because image conditioning constrains the model toward the reference frame. Benchmarked on the same architecture, image-to-video reduced dynamic degree from 39.4 to 32.1, and the researchers described these models as “adhering too closely to the reference image” (arxiv.org, October 2026). Ask for less motion in a shorter take rather than fighting it, or supply the motion directly with a video input.
Which AI video models accept reference images?
Veo 3.1 and Veo 3.1 Fast accept up to 3 reference images (ai.google.dev, September 2026). Seedance 2.0 accepts up to 9 images plus 3 video clips in one request (seed.bytedance.com, September 2026). Kling 3.0 takes up to 3 Elements per clip. Kling 2.6 Pro does not support reference images but offers Motion Control from an uploaded action video (kling.ai, September 2026). Runway Gen-4.5 accepts text and a first frame.
How many attempts does an AI video shot take?
Runway puts text-to-video at 5 to 10 attempts (runway.com, accessed September 2026). Starting from a settled still cuts that sharply, because composition and lighting stop being variables. An attempt as a still costs far less than an 8-second video take. So settle the frame as a still before you pay for motion.
Do I have to disclose AI-generated video in ads?
The rules differ by platform, as of October 2026. TikTok asks advertisers to disclose ads with fully AI-generated or heavily AI-edited images, video or audio, with the AIGC label or their own clear disclaimer. TikTok says it rejects or restricts an ad with undisclosed AI content, and background changes and color adjustments need no label (ads.tiktok.com, October 2026). Meta asks for a disclosure only in ads about social issues, elections or politics. Meta labels other ads itself when it detects signs of third-party AI, such as C2PA metadata (facebook.com/business, October 2026). Google requires election advertisers to disclose synthetic or changed content. For other ads, Google gives advertisers an AI label setting, and says rules in the EU, India and New York require labels on some AI ads (support.google.com, October 2026). This is general information, not legal advice. Verify the current policy before a campaign ships.
Does the input mode fix distorted hands or faces?
No. A correct input settles what is in the frame, not how stably the model redraws it. Hands, faces, and object contact break because each frame is generated without memory of the previous one. Starting from a still where the anatomy is already correct removes much of the risk, but it is a separate fix from the input-mode decision.
Sources
- Veo 3.1 and Veo 3.1 Fast input modes, first and last frame support, and up to three reference images: ai.google.dev, accessed October 2026
- Gemini Omni Flash as Google’s default video model: ai.google.dev, accessed October 2026
- Veo 3.1 preview models listed for shutdown from 22 October 2026 at the earliest: Gemini API deprecations, accessed October 2026
- Sora 2 and Sora 2 Pro removal from the API on 24 September 2026: OpenAI API deprecations, accessed October 2026
- Seedance 2.0 accepting up to 9 images, 3 video clips and 3 audio clips: seed.bytedance.com, accessed October 2026
- Seedance 2.0 first and last frame image-to-video: docs.byteplus.com, accessed September 2026
- Kling 2.6 Motion Control: kling.ai, and Kling 2.6 and Kling 3.0 reference support, including Elements: kling.ai capability map, accessed September 2026
- Runway Gen-4.5 text and image-to-video support and 2 to 10 second durations: docs.dev.runwayml.com, and the 21 January 2026 first-frame changelog entry: runway.com/changelog, accessed September 2026
- Attempt counts and the image-to-video consistency claim: runway.com, published 11 November 2025, accessed October 2026
- Emu Video human preference rates of 70.5% on quality and 63.3% on faithfulness: arXiv:2311.10709, accessed October 2026
- Wan 2.1 same-architecture comparison showing dynamic degree, subject consistency and aesthetic quality changes: arXiv:2506.08456, accessed October 2026
- VBench++ on image alignment limiting temporal consistency: arXiv:2411.13503, accessed October 2026
- Seedance 1.0 technical report on image conditioning and character preservation: arXiv:2506.09113, accessed October 2026
- TikTok ads policy on AI-generated content: ads.tiktok.com, accessed October 2026
- Meta AI info labels on ads: facebook.com/business, and Meta’s social issue, election and political ads rule: transparency.meta.com, accessed October 2026
- Google Ads AI labels and election ad disclosure: support.google.com and Google Ads Help, accessed October 2026
- DesignerBox plans, the video credit range and feature gating: DesignerBox pricing page (designerbox.ai/pricing), September 2026
Model input-mode support verified from provider documentation at ai.google.dev, developers.openai.com, seed.bytedance.com, docs.byteplus.com, kling.ai and docs.dev.runwayml.com as of September 2026. Google’s Veo docs and deprecations page, ByteDance’s Seedance 2.0 launch post, Runway’s tips page and the four arXiv papers were re-checked on 2 October 2026. Research figures from arXiv preprints as cited. Platform policies as of October 2026; this is general information, not legal advice. DesignerBox plans from the DesignerBox pricing page, September 2026. Individual results vary.