Skip to main content
Get started free

AI Video Prompting Guide: 5 Prompt Structures (2026)

Google, OpenAI, ByteDance and Runway each publish a different AI video prompt structure. What each one documents, and how to move a prompt between them.

AI Video Prompting Guide: 5 Prompt Structures (2026)

There is no universal AI video prompt structure. Google documents cinematography first, ByteDance documents camera fourth and caps prompts at 100 words, and OpenAI states that longer prompts reduce the model’s creative range rather than improving it. A prompt tuned for one model degrades when pasted into another, so the fix is to learn the five documented structures and convert between them.

Most prompting guides hide this. They teach one formula, hand you a library of prompts, and leave you to discover that the same prompt returns a clean dolly-in on one model and a shaky zoom on the next. You blame the prompt. You rewrite it four times. The prompt was fine. It was written in the wrong dialect.

Every major provider now publishes its own prompt guide, and those guides do not agree. Two of them contradict the single most repeated piece of prompting advice on the internet.

This guide puts the five structures side by side, names exactly where they disagree, and gives you a conversion routine so one shot brief survives a model switch. It does not hand you 92 prompts. It hands you the rules those prompts were written against, which is the part that transfers.

Key Takeaways

Five providers, five structures. Google, OpenAI, ByteDance, Runway and Kuaishou each publish a different prompt shape. None of them is wrong, and none of them is portable as written.

Camera position is the biggest split. Google’s Veo 3.1 guide opens with cinematography (cloud.google.com, July 2026). ByteDance’s Seedance 2.0 guide puts camera fourth of six (July 2026). Front-loading the camera helps one and fights the other.

Longer is not always better. OpenAI states plainly that “longer, more detailed prompts restrict the model’s creativity” (developers.openai.com, July 2026). Seedance’s guide recommends 60 to 100 words. Neither rewards a 300-word brief.

One camera move per clip. Every guide converges here. Runway’s own guidance is to choose one primary movement per prompt (runway.com, November 2025).

Starting from an image changes the job. With a product photo as the first frame, the image fixes subject, composition, colour and light, so the prompt should describe motion and nothing else (runwayml.com, July 2026).

Video needs Premium. AI video sits on the $75/month tier and up in DesignerBox. Check the plans before you budget a shot list.

What structure an AI video prompt should follow

The structure that matches the model you are rendering on. That is the whole answer, and it is unsatisfying, so here is the practical version.

Every documented structure names the same five decisions: the subject, what it does, where it is, how the camera behaves, and the look. They differ in the order those decisions appear and in how much text each gets. Order matters because these models weight early tokens more heavily. Length matters because two providers explicitly warn that more detail buys less range, not more control.

So a prompt is portable at the level of content and not at the level of form. You can carry the same five decisions across every model. You cannot carry the same paragraph.

The five documented prompt structures, side by side

DesignerBox includes six video models on one bill: Veo 3.1, Veo 3.1 Fast, Sora 2 Pro, Seedance 2.0, Kling 2.6 Pro and Runway Gen-4.5. Five of the providers behind them publish prompt guidance. Here is what each one actually documents.

ModelDocumented structureWhere camera sitsLength guidance
Veo 3.1Cinematography, Subject, Action, Context, Style and ambianceFirstNot specified
Sora 2 ProProse scene description, then separate Cinematography, Actions and Dialogue blocksOwn blockShorter grants more range, longer restricts it
Seedance 2.0Subject, Action, Environment, Camera, Style, ConstraintsFourth of six60 to 100 words
Runway Gen-4.5”The camera [motion] as the subject [action]” for image to videoFirstStart simple, add detail after
Kling 2.6 ProCamera track and subject track kept separateOwn channelNot specified

Sources: cloud.google.com, developers.openai.com, ByteDance’s Seedance 2.0 prompt guide, runwayml.com and fal.ai’s Kling 2.6 motion-control guide, all July 2026.

Read the first two columns together and the design intent shows. Google built a formula for people who think like a director, so the shot comes first. ByteDance built one for people who think like a writer, so the subject comes first and the camera arrives once the scene exists. OpenAI split the two apart entirely and told you to label them.

The parts that are the same everywhere

Underneath the ordering, the guides agree on more than they disagree on:

  • Name one subject and one action. Compound actions chained with “then” or “after that” break in a five-second clip.
  • State the environment concretely. “Abandoned warehouse at dusk” carries information that “dark building” does not.
  • Pick one primary camera move. Two moves in one prompt fight each other.
  • Expect several takes. Runway’s guidance puts professional output at five to ten variations per shot (runway.com, November 2025).

Those four hold on every model in the table. Build your brief on them, then reshape the brief per model. If you want the fuller version of what a brief has to name before ordering ever comes up, the seven layers a realistic prompt specifies covers subject, camera, light, motion, continuity, grade and delivery in detail.

Where the guides disagree, and which side to take

Three disagreements matter enough to change your output. The rest is style.

Disagreement 1: where the camera move goes

Google’s Veo 3.1 guide opens its formula with cinematography, defined as “the camera work and shot composition” (cloud.google.com, July 2026). Runway’s public prompting guide does the same, leading with camera movement (runway.com, November 2025). ByteDance’s Seedance 2.0 guide runs subject, action, environment, then camera (July 2026).

That is a real split, not a phrasing preference. Seedance’s guide names “mixing camera movement with subject movement descriptions” as a top pitfall, which is a structural warning: it wants the subject settled before the camera is introduced, so the model does not attribute the movement to the wrong thing.

Which side to take: follow the model. Lead with the camera on Veo 3.1 and Gen-4.5. Establish the subject first on Seedance 2.0 and Kling 2.6 Pro. On Sora 2 Pro, put the camera in its own labelled block and the question disappears.

Disagreement 2: how long the prompt should be

This is where the internet’s default advice fails. “Add more detail” is repeated everywhere, and OpenAI documents the opposite trade:

“Shorter prompts give the model more creative freedom. Expect surprising results.” and “Longer, more detailed prompts restrict the model’s creativity.” (developers.openai.com, July 2026)

ByteDance recommends 60 to 100 words for Seedance 2.0 (July 2026). Runway’s Gen-4.5 guidance is to start simple, because simple prompts show you how the model interprets motion before you over-constrain the scene (runwayml.com, July 2026).

Which side to take: length is a control dial, not a quality dial. Write short when you are still finding the shot and want variety across takes. Write long when the shot is locked and you need it reproduced. Padding a prompt with adjectives narrows the model’s options without adding instruction, which is the worst of both.

Disagreement 3: how to say what you do not want

Runway’s guide bans negation outright, on the grounds that the model can latch onto the word you are negating: “When you write ‘No camera shake,’ the AI might focus on the word ‘shake’ and give you exactly that” (runway.com, November 2025).

Google’s Veo 3.1 guidance is different. It keeps the exclusion but demands it be concrete, recommending “a desolate landscape with no buildings or roads” over the vaguer “no man-made structures” (cloud.google.com, July 2026).

Which side to take: state the positive first, then the exclusion only if the positive cannot carry it. “Locked-off tripod, no camera shake” gives the model a thing to do and a thing to avoid. “No camera shake” alone gives it only the word shake.

How to write a camera move the model can execute

Camera vocabulary is the part of prompting that looks hardest and is actually the most standardised. The documented sets overlap heavily.

Veo 3.1’s guide lists dolly shot, tracking shot, crane shot, aerial view, slow pan and POV shot for movement, with wide shot, close-up, extreme close-up, low angle and two-shot for composition, and shallow depth of field, wide-angle lens, soft focus, macro lens and deep focus for optics (cloud.google.com, July 2026). Seedance 2.0’s guide names eight movement types: push in, pull out, pan, tracking, orbit, aerial, handheld and locked-off (July 2026).

Those two lists are close enough that one vocabulary covers both. The failure is almost never the word. It is the physics behind it.

The four rules that decide whether a move lands

Give a dolly something to move past. A dolly-in with a flat background reads as a zoom, because parallax is what distinguishes them. fal.ai’s Kling 2.6 motion-control guide is explicit that a dolly needs foreground, midground and background elements to sell (fal.ai, July 2026). Put an object near the lens.

Keep arcs short. The same guide caps a usable orbit at roughly 30 degrees in a five-second clip, because a full orbit at that duration warps geometry (fal.ai, July 2026). If you need a full turn, split it across shots.

Prefer slow. Fast moves produce spatial warping and artifact trails on Kling 2.6, so the guidance is to qualify speed rather than write “fast” bare (fal.ai, July 2026). Seedance’s guide lists an unqualified “fast” as a documented pitfall.

Give the move an ending. “Tracking shot following from the side, then settling back into place” resolves. A move with no endpoint drifts until the clip runs out.

When a move still breaks the subject rather than the frame, the cause is usually elsewhere. Distortion has four separate failure modes and each has a different fix, none of which is a better camera term.

What changes when the shot starts from your product photo

Everything above assumes text to video. Most commercial work does not start there. It starts with a real product photo, and that changes the prompt’s job completely.

Runway’s image-to-video guidance is the clearest statement of the split: the image fixes the subject, the composition, the colour palette and the lighting, so the prompt’s only job is to describe how the frame evolves over the next few seconds (runwayml.com, July 2026). Describing what is already visible in the image wastes the prompt and can fight the reference.

OpenAI documents the same anchoring for Sora 2 Pro. The model “uses the image as an anchor for the first frame, while your text prompt defines what happens next” (developers.openai.com, July 2026). The input image has to match the target video resolution.

So the image-to-video prompt is a motion brief, not a scene description:

Text to videoImage to video
Describe subject, environment, light, camera, styleDescribe motion only
The model invents the productThe photo is the product
Style drift between takesComposition and colour held by the reference
Longer prompt, more decisions to makeShorter prompt, most decisions already made

That is the practical reason a campaign built from one real packshot holds together across a dozen clips while a text-only workflow drifts. Nothing comes out looking generic AI, because frame one is your actual product. Which input mode a shot starts from compares the three options and what each one protects.

Reference handling differs by model, and it is worth knowing before you plan a sequence. Veo 3.1 accepts reference images of a scene, character, object or style to hold an aesthetic across multiple shots, and generates a transition between a supplied start and end frame (cloud.google.com, July 2026). Seedance 2.0 accepts up to nine images, three video clips of two to 15 seconds, and three audio files, capped at 12 files in one generation (July 2026). Kling 2.6 Pro’s Motion Control locks a character’s performance while the camera is directed separately (fal.ai, July 2026).

If a character has to survive the whole sequence rather than one clip, references alone will not do it. Keeping a character consistent across clips covers the anchoring method that does.

Audio is prompted separately, and the syntax is not shared

Audio has its own syntax per model, and getting it wrong is a common cause of a clip that looks right and sounds wrong.

Veo 3.1 documents three formats: quotation marks for dialogue, an SFX label for sound effects, and an ambient noise label for the background soundscape (cloud.google.com, July 2026). Sora 2 Pro wants dialogue in a dedicated block placed below the prose description so the model separates spoken lines from visual description, with speaker labels when more than one character talks (developers.openai.com, July 2026).

Two practical notes. Native audio is the expensive layer in per-second video pricing, so generate it once the visual take is locked rather than on every draft. And most paid social plays silent, so a clip destined for a muted feed does not need it at all.

Convert a prompt from one model to another

Here is the routine. It takes about two minutes per shot and saves the re-rolls.

  1. Strip the prompt to its five decisions. Subject, action, environment, camera, look. Write each as one clause. This is your portable brief.
  2. Reorder for the target. Camera first for Veo 3.1 and Gen-4.5. Subject first for Seedance 2.0 and Kling 2.6 Pro. Labelled blocks for Sora 2 Pro.
  3. Resize to the target’s guidance. Trim to 60 to 100 words for Seedance 2.0. Keep it short on Sora 2 Pro when you want variety across takes.
  4. Rewrite negations as positives. “Locked-off tripod” instead of “no camera movement”.
  5. Check the move against physics. One primary move, an endpoint, depth cues for a dolly, arcs under about 30 degrees.
  6. Move audio into the target’s syntax. Quotation marks and labels for Veo 3.1, a dialogue block for Sora 2 Pro.

Run a cheap draft first. Block the framing, light and motion on a fast model, confirm the shot works, then spend the expensive render once on the take you already trust. Video is priced per second of output, so the cost of a bad prompt scales with the length of the clip you attach to it.

Model-specific starting points sit in the Veo 3.1 prompt library, the Seedance prompt library and the Kling prompt library. Those are written against each model’s own documented structure, so they are a faster starting point than converting by hand.

Where DesignerBox fits

Six video models on one bill removes the reason most teams never test a second model. Switching from Veo 3.1 to Seedance 2.0 is a dropdown, not a new subscription, a new login and a new prompt box to learn. The model catalogue holds 13 models across six providers, six of them video.

The workflow that makes the conversion routine cheap is the one that starts from your product photo. Upload the packshot, animate it, and the prompt only has to describe motion because the image already holds the product. That runs through the Ad Studio for campaign work, or the image to video tool if you want to animate a single photo and see the output before committing to a plan.

Once a prompt is dialled for a model, save the shot as a workflow so the team reruns it for the next product instead of rediscovering the clauses. That is the difference between a prompt library and a production system.

Two honest limits. AI video requires the Premium plan at $75/month or higher, so it is not on Basic or Pro. And video is the most expensive operation in the product, priced per second of output, which is why the draft-cheap-finish-expensive order matters more here than anywhere else. Which models publish enough spec to plan a shot list against scores all six on documentation quality, which is a real selection criterion in a category that changes monthly.

FAQ

What is the best structure for an AI video prompt?

The one the target model documents. Veo 3.1 opens with cinematography, then subject, action, context and style. Seedance 2.0 runs subject, action, environment, camera, style, constraints. Sora 2 Pro uses a prose description with separate cinematography, action and dialogue blocks. The five decisions are the same across all of them. The ordering is not.

How long should an AI video prompt be?

It depends on whether you want control or variety. OpenAI documents that shorter prompts give Sora 2 more creative freedom while longer ones restrict it (developers.openai.com, July 2026). ByteDance recommends 60 to 100 words for Seedance 2.0. Write short while you are still finding the shot, long once you need it reproduced exactly.

Should I use negative prompts in AI video?

Prefer positive phrasing. Runway’s guidance warns that writing “no camera shake” can push the model toward shake, because it latches onto the word (runway.com, November 2025). Google’s Veo guidance allows exclusions but wants them concrete rather than abstract (cloud.google.com, July 2026). Safest pattern: state what you want, then add the exclusion only if the positive cannot carry it.

Can I use the same prompt across different AI video models?

The content transfers, the form does not. Keep the same five decisions, then reorder for the target model, resize to its documented length guidance, and rewrite the audio syntax. Expect the same brief to need three edits to move between models.

Why does my camera movement prompt not work?

Usually physics rather than vocabulary. A dolly with no foreground or background objects reads as a zoom because there is no parallax. An orbit beyond roughly 30 degrees in a five-second clip warps geometry. Fast moves produce artifact trails. And a move with no stated endpoint drifts until the clip runs out (fal.ai, July 2026).

How do I prompt camera movement when starting from a product photo?

Describe motion only. The image already fixes the subject, composition, colour and lighting, so the prompt’s job is how the frame evolves (runwayml.com, July 2026). “Slow push-in, product held centre frame, steam rising” is a complete image-to-video prompt. Re-describing the product in the photo wastes the prompt and can fight the reference.

How many takes should I expect per shot?

Plan for several. Runway’s guidance puts professional output at five to ten variations per shot (runway.com, November 2025). That is the real cost of AI video, which is why blocking the shot on a fast model before committing an expensive render is the standard order of operations.

Sources

All accessed July 2026 unless dated otherwise.

  • Veo 3.1 prompt formula, camera and composition vocabulary, audio syntax, negative-prompt guidance, reference-image and first-and-last-frame behaviour: (cloud.google.com, July 2026)
  • Veo prompt components and camera vocabulary: (deepmind.google, July 2026)
  • Sora 2 prompt template, prompt-length trade-off, dialogue block syntax, and image-as-first-frame anchoring: (developers.openai.com, July 2026)
  • Seedance 2.0 six-step formula, eight camera movement types, 60 to 100 word guidance, documented pitfalls and multimodal input limits: ByteDance Seedance 2.0 prompt guide (July 2026)
  • Runway Gen-4.5 image-to-video guidance on what the image controls versus what the prompt controls, and the start-simple structure: (runwayml.com, July 2026)
  • Camera-first formula, positive phrasing, one primary movement per prompt, and the five-to-ten variations figure: (runway.com, November 2025)
  • Kling 2.6 camera behaviour, orbit and speed limits, dolly depth requirement, motion endpoints and Motion Control: (fal.ai, July 2026)
  • DesignerBox model catalogue, plan pricing and feature gating from the product’s own configuration, July 2026

Prompt structures verified from cloud.google.com, deepmind.google, developers.openai.com, runwayml.com, runway.com, fal.ai and ByteDance’s Seedance 2.0 prompt guide as of July 2026. Provider documentation in this category changes monthly. Individual results vary.

Bogdan

DesignerBox team

Bogdan is part of the team building DesignerBox, the AI creative studio for on-brand campaigns.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Reshoot the clip, not the campaign

Veo 3.1, Sora 2 Pro, Kling 2.6 Pro, Seedance 2.0 and Runway Gen-4.5 are built in. When one model warps a hand, rerun the same shot on another without a second subscription.

Start free

Upload one product photo. Ship the whole campaign, without a photoshoot.