Claude does not render video. It has no video model inside it, and Anthropic’s own help pages say Claude does not generate photos or illustrations either (support.claude.com, accessed August 2026). What a “Claude AI video generator” actually means is Claude connected over MCP to a real video model. Claude writes the brief and calls the tool. Veo, Sora, Kling or Seedance renders the frames.
That distinction sounds pedantic until you budget for it. Search the phrase and every result is a product page promising video inside your chat window. None of them tell you that the model doing the rendering charges per second, that an 8 second clip is the single most expensive thing your account will buy, or that the assistant writing the prompt cannot see what came back.
This guide covers what Claude genuinely contributes to a video, what the connected model contributes, what the whole loop costs in real numbers, and the one failure that no setup guide mentions. Written for marketing leads and founders who want video ads out of a chat window and want to know what they are signing up for first.
Key Takeaways
Claude renders nothing. It plans, writes, and calls tools. Every frame comes from a separate video model reached over MCP.
The split is director and camera. Claude handles script, shot order, pacing and the CTA. The video model handles pixels. Neither does the other’s job well.
Video is priced per second, everywhere. Sora 2 Pro is $0.30 to $0.70 per second of output through OpenAI’s own API depending on resolution (developers.openai.com, accessed August 2026). Credit-based tools charge the same way underneath.
Resolution and duration are coupled. Veo 3.1 only reaches 1080p or 4K at its full 8 second length (ai.google.dev, accessed August 2026). You cannot ask for 4K at 4 seconds.
Claude cannot watch the result. It commissions a clip, gets a file back, and has no way to judge whether the hands are right. Your review step is still a human opening the file.
One photo in beats one prompt in. A text-to-video brief invents a product. Image-to-video from your real packshot keeps the item that ships to the customer.
Can Claude generate video by itself?
No. Claude is a text and reasoning model with vision for reading images you upload. Anthropic’s documentation describes it as processing visual input and generating text and code from it, and Anthropic’s help centre states plainly that Claude does not generate photos or illustrations the way image tools do (support.claude.com, accessed August 2026). There is no video decoder, no diffusion model, and no render step inside Claude.
What changed is not Claude. It is the connection layer.
The Model Context Protocol is an open standard Anthropic published for wiring assistants to outside tools and data (modelcontextprotocol.io, accessed August 2026). A remote MCP server exposes a set of tools. Claude reads their descriptions, decides which one fits your request, fills in the arguments, and calls it. If one of those tools happens to run Veo 3.1, then asking Claude for a product video produces a product video. Claude still did not make it.
What Claude actually contributes
Strip the marketing away and Claude does four jobs in this loop. All four are language jobs.
It turns a sentence into a shot list. You say “a 15 second ad for our ceramic mug aimed at people who work from home”. Claude returns a setting, a camera move, a lighting note, the beat where the product enters frame, and a closing line. That is a brief, and briefs are what video models are bad at receiving from humans.
It writes the prompt in the model’s dialect. Every video model wants different phrasing. Claude has read enough of the documentation to translate your intent into something a specific model responds to, which is otherwise a skill you acquire over months of burned credits. Our AI video prompting guide covers the same craft by hand if you would rather learn it.
It picks the tool. Given 6 video models with different strengths, Claude reads the tool descriptions and routes your brief. Fast draft, or the expensive one with native audio.
It runs the loop. Ask for four variants and Claude issues four calls, tracks the job IDs, and reports back when the renders land. That part is genuinely useful and genuinely boring, which is the best kind of automation.
Notice what is missing from that list. Nothing about pixels. Claude’s contribution ends the moment the tool call fires.
What the video model contributes
Everything you can see. And the model you land on sets hard limits that no amount of clever prompting moves.
| Model | Provider | Native audio | Verified limit |
|---|---|---|---|
| Veo 3.1 | Yes, always on | 4, 6 or 8 seconds. 1080p and 4K only at 8s | |
| Veo 3.1 Fast | Yes | Faster, cheaper draft pass | |
| Sora 2 Pro | OpenAI | Yes, synced | $0.30 to $0.70 per second by resolution |
| Seedance 2.0 | ByteDance | Varies | Strong on motion, cheap per second |
| Kling 2.6 Pro | Kuaishou | Yes | Voiceover, effects and ambience in one pass |
| Runway Gen-4.5 | Runway | Varies | Control-heavy, editorial work |
Veo 3.1 generates audio natively and always on, runs at 24fps, and offers 16:9 or 9:16 (ai.google.dev, accessed August 2026). The duration and resolution coupling in that table is the one people trip over. Asking Claude for “a 4 second 4K clip” gets you an error, not a clip, because that combination does not exist.
Those 6 video models plus 7 image models are the 13 DesignerBox runs, each with a page at designerbox.ai/models if you want the per-model detail before you brief anything.
What the whole loop costs
This is the number every landing page for this keyword leaves out, so here it is straight.
Video is priced by the second of output, not by the request. That is true at the API level and it stays true through any credit system layered on top. At OpenAI’s published rates, an 8 second Sora 2 Pro clip at 1080p costs $5.60 in raw model spend (developers.openai.com, accessed August 2026). Ten variants of that clip is $56 before anyone has picked a winner.
In DesignerBox the unit is credits, and the shape is identical.
| Operation | Credits |
|---|---|
| Generate or edit an image | 5 |
| Create an avatar, 9 images | 25 |
| Brand storyboard | 10 to 25 |
| Video | Credits per second, times duration |
Plans as of August 2026: Free is 112 credits with no card. Basic is $15 a month for 500. Pro is $35 for 1,000. Premium is $75 for 2,500. Ultra is $200 for 8,000. Details on the pricing page.
Now the arithmetic that matters. A Seedance Pro Fast clip at 720p for 5 seconds is 150 credits. Kling Standard at 720p for 5 seconds is 225. Sora 2 at 720p for 8 seconds is 1,600. Veo 3.1 with audio at 8 seconds is 6,400, which is more than a Premium plan holds in a month.
Images are cheap. Video is not. Connecting Claude to the model does not change the physics, it just makes it faster to spend. Budget the credits before you open the chat window, and read what AI video generation costs for the fuller breakdown.
The failure nobody mentions: Claude cannot watch the result
Here is the part the setup guides skip.
Claude can read an image you upload. Anthropic’s vision documentation covers still images. Video is a different input, and Claude has no way to open a rendered MP4 and tell you whether the model gave your hand model six fingers, whether the logo survived, or whether the product on screen is your product.
So the loop is open, not closed. Claude briefs, the model renders, a file lands in your library, and then a person watches it. Every “fully automated video from a chat window” claim quietly depends on you being that person.
Three practical consequences.
Generate a still first, then animate it. Claude can see a still. Ask for the frame, look at it, approve it, then send the approved frame to a video model as an image-to-video input. You have moved the review step to the cheap asset. A rejected 5 credit image beats a rejected 6,400 credit clip.
Draft on a fast model before committing. Run the brief through a cheap fast pass at 720p. Watch it. Only then spend on the expensive render at full resolution.
Ask for fewer, longer variants than you think. Ten 8 second variants means ten files a human has to open. The throughput ceiling on this workflow is your attention, not the model’s queue.
This is the same seam that shows up across agentic AI for content creation: the planning got automated, the judgement did not.
Text-to-video versus your actual product
The second thing that separates a demo from a usable ad.
Ask Claude for “a video of our ceramic mug” with no image attached and the video model invents a ceramic mug. It will be a plausible mug. It will not be yours. The glaze is wrong, the handle is a different shape, and the customer who clicks the ad receives a different object than the one that sold them.
Image-to-video fixes this, and it is the reason the input mode matters more than the model choice for commerce work. You upload the packshot you already have. The model animates that object rather than imagining one. Everything downstream, the on-model shot, the lifestyle scene, the 15 second ad, derives from the same source photo, which is also why the set stays consistent with itself.
DesignerBox is built around that constraint. One product photo goes in and the campaign comes out: product stills, on-model images, video ads, social cuts. The Video Studio is where that runs in the browser. Over MCP, the same operations are tools Claude can call.
How to connect Claude to a video model
Four steps, and the setup is one-time.
- Get an account with credits. Video needs a paid tier in every tool on the market, because per-second pricing makes a free video tier uneconomic. In DesignerBox, AI video sits at Premium and above.
- Add the remote MCP server as a custom connector. In Claude, open Settings, then Connectors, then Add custom connector, and paste the server URL. Anthropic documents the flow for any remote server (support.claude.com, accessed August 2026). The DesignerBox endpoint and the click-through walkthrough are at designerbox.ai/mcp/connect.
- Authorise. Most remote servers use OAuth, so you approve access in a browser window once and Claude holds the session after that.
- Upload your product photo, then brief. Say what the ad is for and who it is aimed at. Let Claude propose the shot list before it spends anything.
Once connected, Claude reaches 43 tools across 8 groups, covering image and video generation, avatars, brand profiles, assets, pipelines and whiteboards. There are also prepackaged Claude skills that bundle a whole job, such as a UGC ad factory, into one instruction. We wrote up the full tool set when DesignerBox MCP launched.
Prompts that survive the render
Two rules cover most of the gap between a good brief and a wasted render.
Say the shot, not the vibe. “Cinematic and premium” gives the model nothing to place in frame. “Slow push in on the mug on a walnut counter, morning window light from the left, steam rising, hold on the glaze at 3 seconds” gives it a shot. Claude is good at expanding the first into the second if you tell it your product and your audience.
Name the beat where the product lands. Ad video fails on timing more often than on quality. Tell Claude which second the product enters and which second the CTA appears, and it will write that into the prompt rather than hoping.
A set of copy-paste starting points lives at designerbox.ai/prompts/veo-prompts. For choosing between models on reliability rather than headline specs, see the most reliable AI video generator.
When the chat window is the wrong place
Honest limits, because this workflow is not universally better.
Precise timing edits. Trimming 400 milliseconds off a cut is a timeline job. Describing it in prose is slower than dragging it.
Anything you need to see while you change it. Colour, crop, and composition tuning want a canvas. Chat is a poor interface for adjustments you judge by eye.
High-volume batch work. For 200 SKUs, a saved workflow beats a conversation. Set the recipe once, run it against the catalogue, and skip the chat entirely.
When you do not have the product photo yet. Shoot or generate the still first. The video step should start from an approved frame.
The chat window wins when the job is a brief: something new, described in words, where the planning is most of the work. It loses when the job is an adjustment.
FAQ
Can Claude generate video by itself?
No. Claude has no video model inside it and Anthropic’s help centre confirms it does not generate photos or illustrations either. Every “Claude video generator” is Claude connected over MCP to a separate video model such as Veo 3.1 or Sora 2 Pro. Claude writes the brief and calls the tool. The connected model renders the frames.
What is MCP and why does it matter for video?
The Model Context Protocol is an open standard from Anthropic for connecting assistants to external tools and data. It matters because it is the only way a chat assistant reaches a video model at all. Without a connector, Claude can describe a video in detail and produce nothing you can upload to an ad account.
How much does generating a video through Claude cost?
The same as generating it anywhere else, because the connector does not change model pricing. Video is billed per second of output. Sora 2 Pro runs $0.30 to $0.70 per second through OpenAI’s API by resolution. In DesignerBox credits, a 720p 5 second Seedance Pro Fast clip is 150 credits and an 8 second Veo 3.1 clip with audio is 6,400.
Will the video show my actual product?
Only if you give it your actual product. A text-only brief makes the model invent a lookalike. Upload your packshot and use image-to-video so the model animates the real object. This is the difference between an ad that matches what ships and one that does not.
Can Claude review the video it generated?
No, and this is the loop’s real limit. Claude reads still images but cannot watch a rendered clip and judge it. Generate a still first, approve that, then animate the approved frame. It moves your review to the cheap asset and keeps a person on the quality call, which is where the judgement has to sit anyway.
Which video model should I pick from Claude?
Let Claude route it, then override when you have a reason. Veo 3.1 for native audio and 1080p or 4K at 8 seconds. Sora 2 Pro for longer synced-audio clips. Seedance 2.0 for cheap motion drafts. Kling 2.6 Pro for voiceover and ambience in one pass. Every model has a page at designerbox.ai/models with the per-model detail.
Do I need a paid plan to generate video from Claude?
Yes. Per-second pricing makes free video tiers uneconomic across the category. In DesignerBox, AI video starts at Premium, $75 a month for 2,500 credits. The free tier’s 112 credits covers image work while you test the connection, which is a sensible way to try the setup before committing.
Sources
- Anthropic, “Can Claude produce images?” (support.claude.com, accessed August 2026)
- Anthropic, “Get started with custom connectors using remote MCP” (support.claude.com, accessed August 2026)
- Model Context Protocol, “Connect to remote MCP servers” (modelcontextprotocol.io, accessed August 2026)
- Google, “Generate videos with Veo 3.1 in the Gemini API” (ai.google.dev, accessed August 2026)
- OpenAI, “Sora 2 Pro model” (developers.openai.com, accessed August 2026)
Model capabilities and pricing verified against provider documentation as of August 2026. Model specs in this category change monthly. Individual results vary.