Limited offer Summer sale, 40% off all annual plans Claim my 40% off
Get started for free

How to Make a Photo Talk With AI for Brand Campaigns

Two ways to make a photo talk with AI, what each costs, which models actually accept a face, and the disclosure rules that land on 2 August 2026.

How to Make a Photo Talk With AI for Brand Campaigns

Two techniques make a photo talk. A talking photo tool takes a portrait plus an audio file and syncs a mouth to it, producing a talking head. An image-to-video model takes the same portrait and generates a moving shot with its own audio. The first suits explainers and training. The second suits ad creative. Cost and legal exposure differ sharply.

Most guides on this topic are written for someone animating a family photo. Yours is a brand asset that will run as a paid ad, and that changes three things: which technique you need, what it costs at volume, and who has to sign what before it ships.

This covers both techniques, which models will actually accept a photo of a face, what a talking clip costs in credits and in subscriptions, and the disclosure rules that become applicable in the EU on 2 August 2026.

Key Takeaways

  • Talking photo and image-to-video are different products. One syncs a mouth to audio you supply. The other generates a new shot with its own audio. Buying the wrong one gets you a stiff talking head in a feed that punishes stiff talking heads.
  • Sora 2 Pro cannot do this. OpenAI’s docs state that “Input images with faces of humans are currently rejected” (developers.openai.com, July 2026). It is the wrong tool for a face, at any price.
  • Seedance 2.0 and Kling 2.6 Pro are the two catalogue paths. Both accept an image and generate native audio including speech (seed.bytedance.com, February 2026; kling.ai release notes, July 2026).
  • Google contradicts itself on Veo 3.1 audio. The Gemini API docs say audio is always on. Google Cloud’s model page lists sound generation as not supported for the same endpoint (July 2026). Test before you plan a campaign around it.
  • EU AI Act Article 50 applies from 2 August 2026. Deployers publishing deepfake content must disclose it. Penalties reach €15M or 3% of worldwide turnover.
  • TikTok bans synthetic endorsements outright. A labelled AI clip of a public figure endorsing a product is still a violation. Labelling does not cure it.
  • The two most-cited AI ad performance stats are fabricated. No Nielsen and DeepMind CTR study exists. No Forrester ROAS report exists. Do not build a business case on either.

What does it mean to make a photo talk with AI?

Making a photo talk means generating video from a single still image in which the subject appears to speak. The AI produces mouth and facial movement that did not exist in the source. Two distinct techniques do this, and they produce different assets for different jobs.

Talking photo. You supply a portrait and an audio file. The model animates the face to match the audio. The head moves, the mouth syncs, the body and background stay largely still. You control the script and the voice completely, because you supplied them.

Image-to-video with generated audio. You supply the same portrait and a text prompt. The model generates a short clip with camera movement, body motion, and its own synthesised audio. You describe what you want rather than supplying it.

Talking photoImage-to-video with audio
You supplyPortrait plus your audio filePortrait plus a text prompt
OutputTalking head, static frameA moving shot with camera work
Voice controlTotal, it is your recordingPrompt-level, the model generates it
Typical lengthMinutes5 to 15 seconds
Best forExplainers, training, localisationAd creative, social, brand film

Both techniques start from a still image, which is what separates them from lip sync. Lip sync re-syncs a mouth in footage you already shot, usually to dub it into another language or fix a line. If you have video already, that is the job you are doing, and when your ad actually needs AI lip sync covers it. This guide is for the case where a photo is all you have.

The mistake is predictable. A team searching for a spokesperson ad finds a talking photo tool, produces a head talking at the camera for thirty seconds, and runs it as a Reel. Vertical social rewards motion and cuts. A static head is the format the placement is worst at.

If your asset is a paid ad, you almost certainly want the second technique. Our comparison of six image-to-video models for ecommerce covers that side in depth for product shots, and the same model behaviour applies to people.

Which AI models can make a photo of a person speak?

Fewer than the marketing suggests. Here is what each provider documents, verified against their own pages in July 2026.

ModelAccepts a face?Native audioVerdict for this job
Seedance 2.0YesYes, including character voiceoversBest documented path
Kling 2.6 ProYesYes, English and ChineseWorks, specs partly unpublished
Veo 3.1YesGoogle’s own pages disagreeTest first
Runway Gen-4.5YesNo audio row in the spec tableMotion only, plan sound separately
Sora 2 ProNo, rejectedYesCannot be used

Sora 2 Pro rejects faces outright

This is the finding that saves the most wasted time. OpenAI’s video generation guide states that “Input images with faces of humans are currently rejected” and that “Real people, including public figures, cannot be generated” (developers.openai.com, July 2026). This is a policy block, not a quality limitation.

OpenAI has also scheduled the Sora 2 API, including sora-2-pro, for shutdown on 24 September 2026 (developers.openai.com, July 2026). Anything you build on it this quarter has a short life. Sora 2 Pro remains excellent for product and scene work where no human face is in the input.

Seedance 2.0 is the best-documented route

ByteDance documents Seedance 2.0 as accepting up to nine images plus audio and video references, generating dual-channel audio across background music, sound effects, and character voiceovers, at up to 15 seconds with multi-shot output (seed.bytedance.com, February 2026).

It also carries a consent requirement worth reading twice: “If you wish to use real human portraits as subject references for video generation, identity verification or prior legal authorization is required” (seed.bytedance.com, February 2026). The model provider is imposing the same rule the law does. Seedance 2.0 sits in the catalogue at up to 15 seconds.

Kling 2.6 Pro handles spoken lines

Kling documents image-to-video with native audio covering speech, dialogue, narration, singing, and ambient effects, in English and Chinese (kling.ai release notes, July 2026). A separate lip sync endpoint exists, though it operates on video already generated in Kling rather than on a raw still.

Clip duration and resolution are not published on a page we could read directly, so treat aggregator figures with caution. Kling 2.6 Pro is the shortest catalogue path when the ad needs someone to say the line.

Veo 3.1 has a documentation conflict

Google’s Gemini API docs describe Veo as natively generating audio with video, always on, with dialogue supported through quoted speech in the prompt. Google Cloud’s model page for the same GA endpoint lists sound generation as not supported (July 2026). DeepMind’s own caveat is the useful one: “creating videos with natural and consistent spoken audio, particularly for shorter speech segments, remains an area of active development” (deepmind.google, July 2026).

Read that as Google telling you speech is the weak spot. Veo 3.1 is outstanding for cinematic motion and 4K output at 4, 6, or 8 seconds. Do not commit a campaign to its dialogue without testing your exact shot first.

The dedicated tools sit outside the catalogue

For a true talking photo, where you upload your own audio file and get a synced mouth, the dedicated category owns it. Runway’s Act-Two accepts an image as a character reference and transfers lip sync and expression from a driving performance video you supply, at up to 30 seconds (help.runwayml.com, July 2026). It is a transfer tool rather than a generator, so you provide both the performance and the voice.

DesignerBox does not include a talking photo tool of that kind. Saying so is more useful than pretending otherwise.

How to make a photo talk with AI, step by step

Audio mixing desk and headphones set up to record the voice track that drives a talking photo
  1. Decide which technique you need first. Script written and voice recorded, and you need it delivered word for word: talking photo. A short spokesperson beat inside an ad: image-to-video. This decision sets the tool, the cost, and the compliance path.
  2. Clear the face before you generate anything. If it is a real person, get a written release that names synthetic generation specifically. If it is a fully synthetic person, confirm they are not readily identifiable as someone real. This step comes before generation, not after.
  3. Prepare the source still. Front-facing, full face visible, even lighting, clean background, highest resolution you have. Every model in this list inherits the flaws in your input and multiplies them across frames.
  4. Draft on the cheap model. Block the shot, the framing, and the beat before spending on a premium render. Throw away most attempts. This is where the credit bill is won or lost.
  5. Review at full size, not on a thumbnail. Mouth sync drifts, teeth deform, and eye lines wander at sizes a thumbnail hides. Check the seconds either side of the audio, where sync usually fails first.
  6. Label it before it ships. Set the platform disclosure toggle, keep the C2PA metadata intact, and record which model produced the asset. The rules below explain why this is not optional.

What does a talking photo cost?

Two costs matter: the per-clip generation cost, and the subscription tier that includes it.

On DesignerBox, video is priced as credits per second times duration, and it is by far the most expensive operation in the product.

Example clipCredits
Seedance Pro Fast, 720p, 5 seconds150
Kling Standard, 720p, 5 seconds225
Sora 2, 720p, 8 seconds1,600
Veo 3 with audio, 8 seconds6,400

Read the bottom row against the plans. Premium is $75 a month for 2,500 credits. One 8-second Veo 3 clip with audio costs 6,400, which is more than two months of that allocation in a single render. AI video requires Premium or higher, and the commercial licence starts at Pro.

The dedicated talking photo tools price by minutes of output instead. Entry plans as published on each vendor’s own pricing page in July 2026:

ToolEntry paid planBilling basis
Magic Hour$15/month, or $10/month billed yearlyCreator
Hedra$15/month, no annual option shownBasic
Synthesia$19/month, or $14/month billed yearlyStarter
HeyGen$29/month, or $24/month billed yearlyCreator
D-ID$4.70/month billed annually ($56/year)Lite

Two cautions on that table. Several of these pages default to the annual view rendered as a monthly figure, so the number you see first is often not what a monthly subscription costs. D-ID’s monthly rate is not published on its pricing page at all, which is why only the annual basis is printed here.

Allowances also differ from the headline. Magic Hour breaks out its talking photo allowance separately from general video, at 42 minutes on Creator against 1.4 hours of general video (magichour.ai, July 2026). Read the allowance for the specific feature you came for, not the plan total. For the wider picture, see what AI video generation costs.

Do you have to disclose an AI talking photo?

In the EU, from 2 August 2026, yes. On TikTok and YouTube, yes today. On Meta, the position is more nuanced than most coverage claims.

EU AI Act Article 50

Article 50 of Regulation (EU) 2024/1689 sets transparency obligations, and it splits them across two parties. Article 50(2) puts a marking duty on the model provider. Article 50(4) is the one that binds a brand: “Deployers of an AI system that generates or manipulates image, audio or video content constituting a deep fake, shall disclose that the content has been artificially generated or manipulated” (artificialintelligenceact.eu, July 2026).

Disclosure must be clear and distinguishable, given no later than the first exposure. Article 50 becomes applicable on 2 August 2026. Systems already on the market before that date have until 2 December 2026 to implement marking, a grace period the Digital Omnibus shortened from six months to three (consilium.europa.eu, May 2026).

Worth separating from the noise: the Omnibus delayed the high-risk rules, not Article 50. Most commentary conflates the two.

Penalties under Article 99(4)(g) reach €15M or 3% of worldwide annual turnover, whichever is higher. A widely repeated figure of €35M or 7% is wrong for this obligation, and applies to prohibited practices under Article 5. The obligation is extraterritorial: a US brand running an EU-targeted campaign is in scope where the output is used in the Union.

Platform policies diverge

PlatformDisclosure required for a commercial ad using AI?
TikTokYes, mandatory. Undisclosed ads are rejected or restricted
YouTubeYes, for realistic and meaningfully altered content
MetaNo general self-disclosure requirement found. Automatic labelling via C2PA detection

TikTok requires creators to label AI-generated or significantly edited content showing realistic people or scenes, and states that undisclosed AI content in ads will be rejected or restricted (tiktok.com and ads.tiktok.com, July 2026).

The sharpest rule in this whole area is also TikTok’s: it does not allow AI-generated content that falsely shows public figures making an endorsement or being endorsed (support.tiktok.com, July 2026). A synthetic celebrity endorsement is banned even when correctly labelled. Disclosure does not cure it.

YouTube requires disclosure when AI makes a real person appear to say or do something they did not, or generates a realistic scene that did not occur. Beauty filters, colour adjustment, background blur, upscaling, and cloning your own voice are explicitly exempt. YouTube also states plainly that disclosing AI content will not limit a video’s audience or its ability to earn money (support.google.com, July 2026).

Meta labels ad content automatically using industry-standard detection including C2PA, rather than requiring self-disclosure for ordinary commercial ads. Self-disclosure is offered as an option in the European Region, California, New York, India, and Taiwan (facebook.com/business, July 2026). Political and social issue ads are a separate and mandatory regime.

Several marketing blogs claim Meta mandates disclosure for photorealistic AI humans in ordinary commercial ads. No Meta-owned page supports that, so do not plan around it.

Detection runs whether you disclose or not

Veo output carries SynthID watermarking and C2PA Content Credentials (ai.google.dev, July 2026). Sora assets carry C2PA. All three major platforms read that metadata. Labels get applied whether or not anyone self-discloses, and on YouTube a metadata-triggered label cannot be removed by the uploader.

The practical question is not only what you must declare. It is what will be detected and labelled regardless.

Whose face can you legally use?

This is where a talking photo differs from every other AI asset. A generated packshot has no likeness exposure. A generated face may have a great deal.

Fully synthetic personDigital replica of a real person
Right of publicity claimNone, no identifiable individualLive exposure
Consent neededNoneWritten, scope-specific release
US state digital replica lawsOut of scopeIn scope
TikTok endorsement banNot triggeredTriggered, labelling does not cure
EU Article 50(4) disclosureStill appliesApplies

That asymmetry is the planning insight. A fully synthetic presenter removes nearly all US likeness exposure and none of the EU disclosure duty.

Right of publicity in the US is state law, with no single federal statute (law.cornell.edu, July 2026). Several states have moved specifically on digital replicas. Tennessee’s ELVIS Act adds voice as a protected right including simulations, effective July 2024.

California’s AB 1836 creates liability for digital replicas of deceased personalities, and AB 2602 makes a contract term permitting digital replica use unenforceable unless it specifically describes the intended uses and the performer had counsel or union representation, both effective January 2025. New York requires written consent for use of a name, portrait, picture, or voice in advertising.

The federal NO FAKES Act is not law. S.4591 was reported out of the Senate Judiciary Committee in June 2026 and placed on the Senate Legislative Calendar, awaiting floor consideration as of July 2026 (congress.gov, July 2026).

Four practical requirements follow for any real person in a talking photo:

  • Get it in writing. New York requires written consent by statute, and written evidence is the only defensible record anywhere else.
  • Name synthetic generation in the release. A photo release drafted before generative video does not obviously authorise a digital replica of the subject.
  • Handle voice separately. Voice is enumerated distinctly under the ELVIS Act and New York law. A likeness release may not cover a cloned voice.
  • Check whether the person is deceased. California protects for 70 years post-mortem, New York for 40.

One caveat cuts against the synthetic route. A generated face that is readily identifiable as a specific real person can still trigger a claim, whatever your input was. The statutory test is identifiability, not provenance.

None of this is legal advice. Any digital replica of a real person, any multi-state campaign, and any reliance on the Article 50 artistic carve-out in commercial work are questions for counsel.

Does AI talking-photo creative actually work?

Start by discarding two numbers you will meet everywhere.

A claimed Nielsen and Google DeepMind study reporting 2.1x higher click-through across 2.3 million impressions does not exist. Neither does a Forrester report claiming 45 to 58% ROAS improvement. Both trace to a single marketing agency page whose own sources list omits them, and both spread through blogs citing each other. No such publications exist at either organisation.

What holds up is more interesting, and less flattering.

IAB’s 2025 Digital Video Ad Spend and Strategy Report found 86% of video ad buyers using or planning to use generative AI for video creative, with roughly 30% of video ad creative built with generative AI and a projected 40% by 2026 (iab.com, July 2025).

Then IAB and Sonata’s The AI Ad Gap Widens, fielded October 2025 to January 2026 across 505 US Gen Z and Millennial consumers and 104 ad executives, found that 82% of ad executives believe those consumers feel positive about AI in advertising, while only 45% actually do. Consumer negative sentiment reached 57%, up 12 points from 2024 (iab.com, January 2026).

A 37-point perception gap, widening, in the same audience most likely to see a talking-photo ad. That is the number to plan around, and almost nobody writes about it.

The disclosure question resolves more gently than the sentiment data suggests. IAB found 73% of consumers say AI disclosure would increase or not change their purchase likelihood, while 89% of advertisers using generative AI disclose at least sometimes and fewer than half always do (iab.com, January 2026). The defensible reading is narrow: disclosure does not appear to hurt. Claims that it lifts trust are not settled, and one controlled academic study found AI-labelled creative was evaluated more negatively than identical human-labelled creative.

So the honest position on a synthetic spokesperson is: cheap to produce, legally intricate, and facing an audience more sceptical than your team assumes. Use it where the format genuinely fits, not because the production cost dropped.

Where DesignerBox fits

DesignerBox is a campaign pipeline, not a talking photo tool. For a portrait plus your own audio file with a synced mouth, the dedicated category does that job better today.

Where it fits is the shot around the spokesperson. Seedance 2.0 and Kling 2.6 Pro both generate speech from a still, and they sit next to the 11 other models in the catalogue on one subscription, in the same workspace as the product stills, the packshots, and the on-model images the same campaign needs. Swap the model per shot without swapping tools. The real cost is not the subscriptions. It is the seams.

Every asset builds from your actual product photo, so nothing comes out looking generic AI. Save the sequence that worked as a workflow and rerun it for the next drop. The Marketing Studio covers the ad side, and Video Ad Composer assembles the finished spot.

For the tools built specifically around synthetic presenters, our breakdown of AI UGC ad tools and what a finished clip costs prices that category against each vendor’s own page, and AI UGC video tools for agencies covers multi-client delivery.

Start free

FAQ

How do you make a photo talk with AI?

Pick your technique first. For a scripted delivery in your own recorded voice, use a talking photo or lip sync tool: upload the portrait, upload the audio, generate. For a short spokesperson beat inside an ad, use an image-to-video model with native audio such as Seedance 2.0 or Kling 2.6 Pro: upload the portrait, describe the shot and the line, generate. Clear the likeness before you generate either one.

Can you make a photo talk with AI for free?

Several dedicated tools offer free tiers, though they generally cap output length, add a watermark, and restrict use to personal rather than commercial purposes. DesignerBox’s free plan includes 112 credits, but AI video requires the Premium tier at $75 a month, and the commercial licence starts at Pro. For an asset that will run as a paid ad, a free tier is rarely a viable path because of the commercial-use restriction rather than the quality.

Which AI model is best for making a photo talk?

Seedance 2.0 has the best-documented path: image input plus native audio including character voiceovers, at up to 15 seconds. Kling 2.6 Pro also generates speech from a still in English and Chinese. Sora 2 Pro cannot be used at all, because OpenAI rejects input images containing human faces. Veo 3.1 excels at motion and resolution, but Google’s own pages disagree on whether it generates audio.

Do I need permission to use someone’s photo in an AI video?

Yes, if the person is real and identifiable. Get a written release that specifically names synthetic generation, because a standard photo release predates the technology and may not cover a digital replica. Handle voice as a separate permission, since several US states protect it distinctly. If the person is deceased, estate consent may still be required for decades. A fully synthetic person avoids most of this.

Do I have to label an AI talking photo as AI-generated?

In the EU, from 2 August 2026, deployers of deepfake content must disclose it under Article 50(4) of the AI Act. TikTok requires disclosure and rejects or restricts undisclosed AI ads. YouTube requires it for realistic altered content depicting real people or events. Meta labels ad content automatically through C2PA detection rather than requiring self-disclosure for ordinary commercial ads. Labels are frequently applied by detection whether or not you declare.

How much does an AI talking photo cost?

On DesignerBox, video is priced as credits per second times duration. A 5-second Kling Standard clip at 720p costs 225 credits, and an 8-second Veo 3 clip with audio costs 6,400, which exceeds the Premium tier’s entire 2,500 monthly allocation. Dedicated talking photo tools price by minutes of output instead, with entry plans from $15 to $29 a month as published in July 2026.

Can I make a photo talk in a language my brand does not have footage for?

Talking photo tools handle this well, since you supply the audio and the model syncs to whatever language it contains. Kling documents native speech in English and Chinese only. For wider language coverage, the dedicated presenter platforms support 30 or more languages. Localisation is the use case where the talking-head format genuinely outperforms an image-to-video clip.

Model capabilities verified against OpenAI, Google DeepMind, Google Cloud, ByteDance Seed, Runway and Kling documentation as of July 2026. Regulatory positions verified against Regulation (EU) 2024/1689, Council of the EU, congress.gov and platform policy pages as of July 2026. Competitor pricing verified against each vendor’s own pricing page in July 2026, with billing period noted. Nothing here is legal advice. Model specifications and platform policies in this category change frequently. Individual results vary.

Vytas

Founder at DesignerBox

Vytas is a founder at DesignerBox, from the team behind LoadFocus, FocusBox and PostNext. He writes about turning one product photo into a full campaign, and the pipelines that keep every asset on brand.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Every top video model, one bill

Veo 3.1, Sora 2 Pro, Kling 2.6 Pro, Seedance 2.0 and Runway Gen-4.5 are built in. Switch models per shot without a second subscription or a second login.

Start free

Upload one product photo. Ship the whole campaign, without a photoshoot.