Limited offer Summer sale, 40% off all annual plans Claim my 40% off
Get started for free

Turn One Product Photo Into a Video Ad, Shot by Shot

One image makes one shot, not an ad. Turn a product photo into a shot list of stills first, then animate each one. The pipeline, costs and label rules.

Turn One Product Photo Into a Video Ad, Shot by Shot

You cannot turn one product photo into a video ad in one step. Image-to-video models animate the frame you give them and return a single continuous shot of 8 to 20 seconds. An ad is three to five shots. So the working pipeline has two hops: generate a shot list of stills from your one photo first, then animate each still separately, then cut them together.

Most guides skip the first hop. They tell you to upload the photo, write a motion prompt, and generate. What comes back is a slow push-in on the same composition you started with. Generate five of those and you have five slow push-ins on the same composition, which cuts together into one long shot that happens to change speed.

This guide covers the pipeline that actually produces an ad: what to check in the source photo, how to derive a shot list of stills from it, how to animate each one, what the whole thing costs in credits, and the disclosure rules that decide whether the finished ad runs.

Key Takeaways

  • One image makes one shot. Veo 3.1 generates 8-second sequences (deepmind.google, July 2026). Sora 2 Pro supports 4, 8, 12, 16 and 20-second clips (developers.openai.com, July 2026). None of those is an ad on its own.
  • The missing step is stills, not prompts. Derive three to five still frames from your one photo first, then animate each. Stills cost 5 credits each, so the step that decides the ad is the cheapest one in it.
  • Your source photo caps everything downstream. Motion cannot add detail that is not in the frame. A soft, cropped, or low-resolution original produces a soft, cropped, low-resolution clip.
  • The cost lever is the model, not the length. A 15-second three-shot ad runs about 450 credits on the cheapest video rate and about 12,000 on Veo 3 with audio. Same ad, 26x the price.
  • AI video starts at Premium. Video generation needs the $75/month tier or higher. Basic and Pro do not include it.
  • Meta mostly labels for you. TikTok does not. Meta applies its own AI info label through automated detection (about.fb.com, July 2026). TikTok puts the duty on the advertiser, and duplicating a campaign resets the toggle (ads.tiktok.com, July 2026).

Why one image makes a shot, not an ad

A video model reads your still frame and predicts motion forward from it. It holds one camera, one composition, one lighting setup, because that is all the frame gave it. What comes back is a shot.

The published limits make this concrete. Google generates Veo 3.1 in 8-second sequences and outputs at 1080p or 4K, with audio generated natively (deepmind.google, July 2026). OpenAI’s sora-2-pro accepts clip lengths of 4, 8, 12, 16 and 20 seconds, and extends up to a total of 120 seconds (developers.openai.com, July 2026). Those are shot lengths.

A product ad that holds attention cuts. It opens on the product in context, moves to a detail, shows the product in use or on a model, and lands on the offer. Four beats, four compositions, four camera positions. One frame cannot supply four camera positions, no matter how the motion prompt is worded.

So the question changes. It stops being “what prompt turns this photo into an ad” and becomes “what four photos do I need, and how do I get them from the one I have”.

Step 1: Check the photo before you animate it

Motion is generative, detail is not. The model invents plausible movement, but it cannot invent the label text it never saw or the stitching that was already blurred. Every flaw in the source gets carried into the clip and then held on screen for eight seconds instead of glanced at in a feed.

Check four things before you spend anything:

CheckWhat passesWhy it matters
Resolution1080px on the short edge, minimumThe clip inherits the source detail. Upscaling after does not recover it
Product edgesClean separation from the backgroundSoft edges become warping artifacts once the camera moves
Label and logo legibilityReadable at 100%Text is what models distort first, and what buyers check
CropProduct fully in frame with marginThe model cannot extend what was cropped out

If the photo fails on resolution, decide whether to fix it or reshoot before you build on it. We covered that trade-off in what to do with low resolution product images. If it passes, the rest of the pipeline is worth running.

One more check that costs nothing: look at the photo and ask whether it is your product or a product like yours. Generated backgrounds and generated scenes are fine. A generated product is a different item that resembles yours, and buyers notice on delivery. Product photo accuracy covers what to verify before shipping.

Step 2: Turn one photo into a shot list of stills

This is the hop the category skips, and it is the one that makes the output read as an ad.

Write your four beats first, as compositions rather than as copy. A workable default for a product ad:

  1. Hero. The product in a clean or styled context. Establishes what it is.
  2. Detail. A close crop on the material, the texture, the mechanism, the finish. Proves quality.
  3. Context. The product in use, on a model, or in the room it belongs in. Shows the outcome.
  4. Offer. The product framed with space for the price, the claim, or the CTA overlay.

Now generate each beat as a still image from your one source photo. Angle generation gives you the hero and the detail from the same product, since Photo Angles is built to produce every camera angle from a single photo. Scene generation gives you the context beat. The offer beat is usually the hero shot recomposed with negative space.

Two reasons this order is right. Stills are cheap, at 5 credits per image, so you iterate on composition at almost no cost before committing to video rates that run hundreds of credits per clip. And a still is reviewable in a second, while a clip takes a minute to generate and thirty seconds to watch.

Approve the four stills as a set. If the product looks like the same item in all four, with the same colour and the same finish, the ad will cut together. If it drifts between frames, fix it here, because motion will not hide it. Our guide to keeping characters and products consistent across clips covers what causes the drift.

For the shot list itself, the complete product photography shot list is a longer reference on which angles earn their place.

Step 3: Animate each still, one shot at a time

Now you animate. One still in, one shot out, four times.

Keep the motion small. The frame already contains the composition you approved, and the job of the motion is to make it live rather than to reinvent it. Push-in, slow orbit, rack focus, a hand entering frame. Camera moves that stay inside the geometry the still established will hold. Moves that ask the model to reveal something outside the frame will invent it, and what it invents is the part that looks generic AI.

Name four things in each motion prompt: the camera move, the speed, the light behaviour, and what stays still. The last one does more work than the other three, because it tells the model which pixels are the product. How to write realistic AI video prompts breaks the full structure down layer by layer.

Model choice is per shot, not per campaign. A detail shot on a static product wants a model that holds texture. A context shot with a person moving wants a model that handles bodies. The six video models compared for ecommerce covers which holds what, and DesignerBox exposes all six in its model library so you can switch between shots without switching tools or bills.

Generate each shot once at low cost first. Watch it. Only re-generate at a higher-priced model when the composition is right and you want the fidelity.

Step 4: Assemble, then cut for the placement

Four 5-second shots is 20 seconds of footage. That is your master, not your ad.

Cut the master to the placement. A feed ad and a Reels ad want different lengths and different first frames, and the first frame is the one that decides whether the rest gets watched. Lead with whichever of your four shots survives being seen at thumbnail size with the sound off.

Aspect ratio is a crop decision, and it is why the margin in Step 1 mattered. A 9:16 crop of a 1:1 composition loses the sides. If you generated the stills with margin, you crop. If you did not, you regenerate.

Assembly is also where the ad gets its copy, its price, and its CTA, none of which the video model produces. Video Ad Composer handles the assembly step from footage and images, and Marketing Studio is where the whole sequence lives if you are running it repeatedly rather than once.

For platform-specific specs, the TikTok video ad walkthrough covers the published dimensions, the safe zone, and the length limits, including which widely quoted numbers are not actually TikTok’s.

What a three-shot ad actually costs

Video is priced per second of output, which makes the model the cost lever and the length a secondary one. DesignerBox publishes four reference points: Seedance Pro Fast at 720p costs 150 credits for 5 seconds, Kling Standard at 720p costs 225 for 5 seconds, Sora 2 at 720p costs 1,600 for 8 seconds, and Veo 3 with audio costs 6,400 for 8 seconds.

Divide those out and you get the per-second rate, which is what to budget against. Here is the same 15-second ad, three shots of 5 seconds each, priced across all four:

ModelCredits per second15-second adAgainst a 2,500-credit month
Seedance Pro Fast, 720p30450About 5 ads
Kling Standard, 720p45675About 3 ads
Sora 2, 720p2003,000One ad exceeds the month
Veo 3 with audio80012,000One ad exceeds Ultra’s 8,000

Same ad, same length, 26x between the cheapest and the most expensive. That spread is the single biggest decision in the pipeline, and it is made per shot rather than per campaign.

The stills, meanwhile, cost 15 credits for all three at 5 credits each. The step that determines whether the ad works is roughly 3% of the cost of the step that renders it. Spend your iterations there.

Two limits worth stating plainly. AI video generation requires the Premium tier at $75 a month or higher, so Basic and Pro do not include it. And Premium’s 2,500 monthly credits do not cover a single 15-second ad at Sora 2 rates. Budget video separately from images, and check the full cost breakdown by model before committing a campaign to a rate. The still-image steps that decide whether the ad works run on the free plan, so the cheap half of this pipeline costs nothing to test.

The label rules that decide whether it runs

A widely repeated claim says Meta now requires every advertiser to disclose AI-generated content. That is not what Meta’s published policies describe, and building a process around it wastes effort on the wrong platform.

What Meta documents is narrower. It applies its own AI info label to ads created or significantly edited with its generative tools, and from 1 June 2026 it uses automated detection to identify media made with third-party generative tools and label those itself, with no advertiser action required (about.fb.com, July 2026). Advertiser self-disclosure is required for ads about social issues, elections, or politics. For a product ad, the labeling is largely something the platform does.

TikTok is the opposite, and it is where the process work belongs. AI-generated, synthetic, or manipulated media is a mandatory disclaimer category, covering both fully generated media and real source material significantly modified by AI. Undisclosed AI content gets the ad rejected or restricted. The mechanism is a self-disclosure toggle in Ads Manager reading “This ad contains AI-generated content”, and creating or duplicating a campaign resets that toggle even when you are reusing creative you already marked (ads.tiktok.com, July 2026). That reset is what catches people scaling a winner.

The industry standard points the same way. The IAB published its AI Transparency and Disclosure Framework on 15 January 2026, and it is deliberately risk-based instead of universal: disclosure is required only when AI materially affects authenticity, identity, or representation in ways that could mislead. Routine production work, background AI tooling, and clearly stylised creative do not trigger it (iab.com, July 2026).

An animated product photo of a product you actually sell sits at the low-risk end of that framework. A synthetic person presenting it does not. Know which one you built.

FAQ

Can you turn a product photo into a video ad with one prompt?

You can turn it into one shot with one prompt. A shot is not an ad. Image-to-video models return a single continuous take from the frame you supply, and published limits sit at 8 seconds for Veo 3.1 and up to 20 seconds per clip for Sora 2 Pro (deepmind.google and developers.openai.com, July 2026). An ad that holds attention cuts between three to five compositions, which means three to five source frames.

How many shots does a product video ad need?

Three to five works for most placements. A workable default is hero, detail, context, and offer: what the product is, why it is well made, what it does for the buyer, and what you want them to do. Fewer than three tends to read as a moving photo. More than five inside 15 seconds tends to read as rushed.

What resolution does the source product photo need to be?

At least 1080px on the short edge, with the product fully in frame and clean edges against the background. Motion is generative but detail is not, so the model cannot add sharpness the original never had. Whatever softness exists in the still gets held on screen for the length of the clip.

Why do AI product videos look generic?

Usually because the motion prompt asked the model to reveal something outside the source frame. Inside the frame, the model has real pixels to work from. Outside it, the model invents, and the invented part is what reads as generic AI. Keep camera moves within the geometry the still already established, and name explicitly what should stay still.

How much does a 15-second AI video ad cost?

It depends almost entirely on the model. Priced per second of output, a 15-second three-shot ad runs about 450 credits at the cheapest rate and about 12,000 credits on Veo 3 with audio, a 26x spread on identical footage length. AI video generation requires the Premium tier at $75 a month or higher.

Do you have to label an AI product video ad?

On TikTok, yes. AI-generated or significantly AI-modified media is a mandatory disclaimer category, and undisclosed content gets rejected or restricted. Use the “This ad contains AI-generated content” toggle, and re-check it after duplicating a campaign, because duplication resets it (ads.tiktok.com, July 2026). On Meta, the platform applies its own AI info label through automated detection and no advertiser action is required outside the social issues, elections and politics category (about.fb.com, July 2026).

Is it cheaper to generate stills first or go straight to video?

Stills first, by a wide margin. An image costs 5 credits, so a four-frame shot list costs 20 credits to build and iterate. Video rates start around 30 credits per second and reach 800. Getting the composition wrong in stills costs almost nothing. Getting it wrong in video costs the whole clip.

Can one product photo produce a whole campaign?

The still set can, which is the point of generating it. Once you have four approved compositions of the same product, they feed video shots, static ads, PDP images, and social posts from one source. That is the argument for treating the stills as the asset and the clips as one output of it rather than the goal.

Model clip limits verified from deepmind.google and developers.openai.com, platform disclosure rules from about.fb.com, ads.tiktok.com and iab.com, as of July 2026. DesignerBox pricing and credit costs from the product’s own published rates. Individual results vary.

Cristian

Head of Content at DesignerBox

Cristian covers AI product photography, video ad tools and model comparisons. He runs the same prompt and the same product across models, then publishes the output side by side, so you pick on evidence instead of marketing copy.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Every top video model, one bill

Veo 3.1, Sora 2 Pro, Kling 2.6 Pro, Seedance 2.0 and Runway Gen-4.5 are built in. Switch models per shot without a second subscription or a second login.

Start free

Upload one product photo. Ship the whole campaign, without a photoshoot.