Generating one AI clip takes about 30 seconds to 2 minutes. Getting a finished, on-brand video takes longer, because a clip runs 4 to 20 seconds and an ad runs 30. The time goes into re-rolls, keeping the product consistent between shots, and the stitch. Front-load those decisions into cheap stills and the whole job compresses.
The fast numbers you see quoted are real, and they measure the wrong thing. Runway’s own beginner guide puts a professional video at 15 minutes and notes that most creators generate three to five versions before they get their result (runway.com, November 2025). Both statements are true. Read together, they describe a 15-minute job that runs three to five times, and that is before anyone stitches the clips or checks whether the product looks like the same product in shot two.
So the useful question is not how fast a model generates. Every model in the catalog generates faster than you can review the output. The useful question is how many generations stand between a brief and something you would put spend behind.
This guide covers where the minutes actually go, the order of operations that cuts them, what each model’s clip cap does to your edit, and what speed costs when you pay per second of output.
Key Takeaways
- Generation is already the fast part. A clip returns in under two minutes. The queue between you and a finished video is re-rolls, not render time.
- Count generations, not minutes. Three to five versions per shot is the published norm (runway.com, November 2025). A four-shot ad at that rate is 12 to 20 generations.
- Clip caps decide your edit. Veo 3.1 generates 4, 6, or 8-second clips (ai.google.dev, July 2026). Sora 2 Pro goes up to 20 seconds (platform.openai.com, July 2026). A 30-second spot is always multiple generations.
- Stills are the cheapest place to be wrong. An image costs 5 credits. Video bills per second of output, so approving composition as a still first is the single biggest time saver in the pipeline.
- Re-rolls re-bill. Video is the most expensive operation in the product. A Veo 3 clip with audio at 8 seconds costs 6,400 credits, and the second attempt costs the same as the first.
- Speed compounds only when the setup is reusable. The first campaign is slow because you are making decisions. The second one is fast because the decisions are saved.
- AI video needs the Premium tier. Video generation starts at $75/month. Basic and Pro do not include it.
How long does it actually take to make an AI video?
One clip: 30 seconds to 2 minutes of generation, plus the time to write the prompt and watch the result. One finished 30-second video: closer to half a day the first time you make it, and under an hour once your references, shot list, and prompt patterns are saved. The gap between those two numbers is re-rolls and continuity work, not compute.
That range surprises people because the demo and the deliverable get quoted with the same word. A demo is one clip generated once from a prompt someone already knows works. A deliverable is four shots that have to look like the same product, cut together, hold a brand, and survive a review.
The fast number and the real number measure different things
There are four clocks running on any AI video job, and guides almost always optimize the smallest one.
| Clock | What it covers | Typical share of the job |
|---|---|---|
| Setup | Brief, references, shot list, aspect ratio | Large on job one, near zero after |
| Generation | The model producing a clip | Under 2 minutes per clip |
| Re-roll | Regenerating shots that missed | The biggest single block |
| Assembly | Stitch, audio, captions, format variants | Fixed, and predictable |
Generation is the clock everyone races. It is also the one clock you cannot meaningfully improve, because a faster model returns a clip you still have to watch, judge, and usually run again.
Re-rolls are where the time actually sits, and they are the clock you can move. Every re-roll is a decision that was not made before generating. Ambiguous prompt, unclear composition, a product reference the model never received. Fix those upstream and the re-roll count drops without anything getting faster.
Where the time actually goes
Three specific failures produce most re-rolls, and each has a cheap fix.
The composition was never decided. Prompting a motion description and hoping a good frame comes back means judging composition and motion at the same time, on an asset that bills per second. Decide the frame first, as a still.
The product drifts between shots. Shot one and shot three come back as recognisably different objects. This is the most expensive failure because you only see it at the edit, after every clip is paid for. Our guide to keeping a character or product consistent across clips covers what causes drift and which reference inputs stop it.
The clip length did not match the beat. Generating 8 seconds for a 3-second cut wastes 5 seconds of paid output. Length is a prompt-time decision on most models, and getting it wrong is pure waste.
None of these are model problems. All three are order-of-operations problems.
Generate stills before you generate video
This is the fastest single change available, and it is counterintuitive because it adds a step.
An image costs 5 credits and returns in seconds. Video bills per second of output and takes a minute or more. So you should be making every composition, lighting, angle, and product-accuracy decision in the cheap medium, then animating only frames you have already approved.
The working order:
- Write the beats, not the copy. Four compositions for a typical ad: hero, detail, context, offer.
- Generate each beat as a still. From your real product photo, so the object in frame is your object.
- Approve the stills as a set. Same colour, same finish, same item in all four. Fix drift here where it costs 5 credits, not at video rates.
- Animate each approved still. Image to video takes a frame you already trust and adds motion to it, which is a much narrower job than inventing a scene.
- Stitch, then cut format variants.
Steps 1 to 3 feel like a detour. They are the reason step 4 lands on the first or second attempt instead of the fifth. Turning one product photo into a video ad shot by shot walks the same pipeline in more detail, and script to storyboard covers the version that starts from a script rather than a photo.
How to cut re-rolls in half
Re-rolls fall when the model has less to guess. Four levers, in order of effect.
Give it a reference, not a description. A product photo, a character reference, or an approved still removes an entire category of guessing. Text alone asks the model to invent the subject, and it will invent a slightly different one each run.
Write shorter, more specific prompts. Runway’s guidance puts the useful range at 15 to 30 words (runway.com, November 2025). Past that, a prompt reads as a list of competing instructions and the model satisfies some of them. Our breakdown of prompt structure for realistic AI video covers what to include and what to cut, and the Veo prompt patterns page has copy-paste starting points.
Pick a model that repeats. Some models return a usably similar result on the same prompt, and some vary widely. Variance is a re-roll multiplier. Six models ranked on what you can plan around covers which ones hold a shot.
Cap the attempts. After five or six similar tries, the prompt is not the problem. Change the input frame, the model, or the beat instead. Runway recommends the same rethink at that point (runway.com, November 2025), and it is the cheapest rule in the list because it caps the downside on any single shot.
Clip caps decide how many generations your video needs
No current model generates a 30-second multi-shot ad in one pass. Published limits, verified this month:
| Model | Clip length | Notes |
|---|---|---|
| Veo 3.1 | 4, 6, or 8 seconds | 720p, 1080p, or 4K, with natively generated audio (ai.google.dev, July 2026) |
| Sora 2 Pro | Up to 20 seconds | 1920x1080 or 1080x1920 among supported resolutions (platform.openai.com, July 2026) |
| Runway Gen-4.5 | Text to video and image to video | Billed at 12 credits per second on Runway’s own plans (help.runwayml.com, July 2026) |
Read that as a planning constraint. A 30-second spot is four to six generations minimum, plus a stitch, whatever the marketing says about speed. Budget the edit, not just the render.
It also explains why “make it longer” is the wrong instruction. Longer clips hold one camera position, so a 20-second single shot reads as a slow zoom rather than an ad. Cuts come from separate generations. The full catalog and its per-model behaviour sits at the model listing, and Runway’s model page is at Runway Gen-4.5.
A worked example: one 30-second product ad
Numbers make the trade obvious. Here is the same four-shot ad produced two ways, counted in generations rather than minutes.
| Step | Prompt-first | Stills-first |
|---|---|---|
| Compositions decided | In video, per attempt | 4 stills, 20 credits total |
| Video generations needed | 12 to 20 | 5 to 7 |
| Where mistakes surface | At the edit, after paying | At the still, before paying |
| Continuity check | Shot by shot, late | Once, on the approved set |
| Re-roll exposure | Every miss bills at video rate | Most misses bill at 5 credits |
The stills-first column is not faster per clip. Each generation takes the same 30 seconds to 2 minutes. It is faster overall because it needs roughly a third fewer video generations, and it moves the failures into the medium where being wrong is cheap.
The second time you run this ad for a different product, the setup column disappears entirely. That is the real speed gain, and it only arrives once something is saved.
What speed costs in credits
Video is by far the most expensive operation, and it bills per second of output. That makes the re-roll count a budget line, not a workflow detail.
| Operation | Credits |
|---|---|
| Generate or edit an image | 5 |
| Video | Credits per second, by model, times duration |
| Veo 3 with audio, 8 seconds | 6,400 |
A single Veo 3 clip with audio costs more than the Premium tier’s entire 2,500 monthly allocation. Run it three times and the arithmetic stops being a rounding error. This is the honest reason to approve compositions as stills: at 5 credits, being wrong is free, and every wrong decision you catch there is a four-figure decision you do not make later.
Plan allocations are 112 credits free, 500 on Basic, 1,000 on Pro, 2,500 on Premium, and 8,000 on Ultra. Video generation starts at Premium. What each plan buys in clips does the arithmetic per tier, and the 2026 AI creative cost benchmark puts those numbers against production costs.
When fast is the wrong target
Speed pays on volume work: ad variants, format cuts, seasonal refreshes, listing videos. Those jobs have a known-good template and the only variable is throughput.
Speed does not pay on the first version of anything. A launch film, a new brand look, or a hero spot needs decisions made slowly and then locked, because everything downstream inherits them. Rushing the first one guarantees you rush all fifty that copy it.
The practical split: spend the time on the setup, then never spend it again. Once a campaign works, save it as a rerunnable campaign variant workflow so the next product runs the same sequence without the same decisions. That is where AI video actually gets fast, and it is a second-campaign benefit, not a first-clip one.
For a shortcut on the most common job, the product video generator runs the photo-to-video path with the setup already made.
FAQ
How long does it take to make an AI video?
A single clip generates in about 30 seconds to 2 minutes. A finished 30-second video with multiple shots takes a few hours the first time, including re-rolls and assembly, and under an hour once your references and prompts are reusable. The variable is the number of generations, not the speed of any one of them.
What is the fastest way to make an AI video?
Start from an image, not a text prompt. Generate your compositions as stills at 5 credits each, approve them, then animate the approved frames. This removes most re-rolls because the model is adding motion to a decided frame rather than inventing the subject and the motion at the same time.
Why does my AI video need so many re-rolls?
Usually because the prompt is doing two jobs at once, describing both what is in frame and how it moves. Split them: fix the frame as a still, then prompt only the motion. Missing reference images are the other common cause, since text alone lets the model reinvent the subject on every run.
Can AI generate a 30-second video in one go?
Not as a multi-shot ad. Veo 3.1 generates 4, 6, or 8-second clips (ai.google.dev, July 2026) and Sora 2 Pro reaches 20 seconds (platform.openai.com, July 2026). Longer single generations hold one camera position, so cuts have to come from separate clips stitched together.
Is it faster to start from text or from an image?
From an image, for anything involving a real product. Text to video invents the subject, which means judging the subject and the motion on every attempt. Image to video fixes the subject first, so re-rolls only test motion. Text to video is faster only when the subject genuinely does not matter.
Does generating at a lower resolution make it faster?
Somewhat, and it is a reasonable way to test a shot before committing. Generate a draft to check composition and motion, then rerun the approved version at delivery resolution. Treat the draft as a preview, since the final pass still bills at full rate.
How many credits does a re-roll cost?
The same as the original. There is no discount for a second attempt, which is why re-roll count is the number worth managing. On an 8-second Veo 3 clip with audio at 6,400 credits, three attempts costs 19,200, more than the Premium tier’s monthly allocation seven times over.
Model capabilities and clip limits verified against Google, OpenAI and Runway documentation as of July 2026. Production timings are ranges, not guarantees. Individual results vary.