Skip to main content
Scale your content with AI and keep your brand, now from Claude, ChatGPT and Cursor. DesignerBox in your AI chat Start DesignerBox MCP

AI Voice Cloning for Brand Video: How It Works (2026)

AI voice cloning copies one voice from 1 to 2 minutes of audio. See how it works, 4 things that keep a brand voice the same, and the consent and label rules.

AI Voice Cloning for Brand Video: How It Works (2026)

AI voice cloning makes a digital copy of one person’s voice from a recording. You give the tool a sample, the tool learns how that person sounds, and it then reads any script in that voice. For brand video, a clone gives every ad the same narrator. It also needs the speaker’s consent, and some ad platforms ask for a label.

A brand that makes video every week has a voice problem. The first ad has one narrator. The tenth ad has a narrator who sounds a little faster and a little brighter.

This guide covers the two kinds of clone, the choice between a preset voice and a cloned voice, and four things that keep one voice the same in every video. It then covers consent law and the label rules on TikTok, Meta and YouTube.

Key Takeaways

  • AI voice cloning copies one real voice. The tool learns the voice from a recording and reads new scripts in it.
  • There are two kinds of clone. An instant clone needs about 1 to 2 minutes of audio. A professional clone needs at least 30 minutes.
  • A preset voice avoids the consent step. Use a clone only when the voice itself is known, such as a founder’s voice.
  • Four things keep a voice the same. The same voice ID, the same settings, the same script style and the same audio finishing.
  • Consent comes first. New York asks for written consent and California asks for prior consent before a voice is used in advertising.
  • Label rules differ by platform. TikTok names voice cloning in its ads policy. YouTube does not ask for a label when you clone your own voice.

What is AI voice cloning?

AI voice cloning is the use of AI to copy how one person speaks. The copy keeps the pitch, the accent, the pace and the small habits of the real voice. After that, the tool turns typed text into speech in that voice. This is different from a preset voice, which belongs to the tool’s library and to no real customer.

A clone copies everything in the sample, good and bad. ElevenLabs writes that the AI “will attempt to mimic everything it hears in the audio”, including the speed, the accent and the breathing (elevenlabs.io/docs, October 2026). A noisy room in the sample becomes a noisy voice in every ad.

Voice cloning sits next to three other audio jobs. AI voiceover for UGC ads covers where the voice in an ad comes from. AI dubbing covers the same ad in a new language. This guide covers the copy of one voice and the rules around it.

How does AI voice cloning work?

Voice cloning works in three steps.

  1. Record a sample. One person reads in a quiet room with one microphone. The reading style should match the style the ads need.
  2. Make the voice. The tool reads the sample and saves a voice under a name or an ID.
  3. Read a script. You type the script, choose the saved voice and make the audio file.

Vendors offer two kinds of clone. ElevenLabs is a clear example because it documents both.

Instant cloneProfessional clone
Sample lengthAbout 1 to 2 minutesAt least 30 minutes, with 2 to 3 hours as the best case
How it is madeThe tool reads the sample onceA model is trained on the sample
Time to first useShortUsually 3 to 6 hours of training
Whose voiceA voice you have the right and consent to cloneYour own voice only, after a verification step
FitsA test, or a short seriesA voice the brand will use for years

Every cell in the table comes from the ElevenLabs documentation (elevenlabs.io/docs, October 2026).

Woman with curly hair holds a microphone at an event under blue lights, like a speaker whose voice a brand records as a sample

Other vendors set other sample sizes. Microsoft’s personal voice makes a voice from a speech sample of 5 to 90 seconds plus a recorded statement from the speaker (learn.microsoft.com, October 2026). Higgsfield describes a custom voice made from a recording of up to two minutes or an uploaded audio file (higgsfield.ai, October 2026).

Preset voice or cloned voice

Most brand video does not need a clone. A preset voice is a stock voice from a tool’s library. You choose it once and use it in every video. No person has to sign anything, because the vendor already holds the rights to that voice.

Woman in a yellow coat with silver headphones around her neck in front of a round light, like a narrator a brand hires for its videos

A cloned voice is worth the extra work in three cases:

  • The founder already speaks for the brand. Viewers know the voice, and a preset would sound like a stranger.
  • A hired voice actor is the brand’s narrator. The clone lets the actor approve scripts and skip the recording booth.
  • A series has one known host. The clone keeps the host’s voice when the host is not free to record.

In each case the person is real, so the consent rules below apply. AI spokesperson covers the same choice for the face: a stock avatar, a custom character or a clone of a real person.

Four things that keep one voice the same

A saved voice does not sound the same in every video on its own. ElevenLabs says its settings are “nondeterministic”, so the same voice, settings and model give “slightly different output” each time (elevenlabs.io/docs, October 2026). Four things reduce that difference.

Four tiles, one per thing that keeps an AI voice the same in every video: the same voice ID, the same settings, the same script style and the same audio finishing.

1. The same voice ID. Use one saved voice for one narrator. Write the voice name or ID in the brand document, next to the logo and the colors.

2. The same settings. Save the numbers and reuse them. In ElevenLabs the settings are stability, similarity, style exaggeration and speed. Speed runs from 0.7 to 1.2, and 1.0 is the default. A higher stability setting gives a steadier read with less emotion. Also keep the same speech model, because a new model reads the same voice in a new way.

3. The same script style. The script changes the sound. Keep sentence length and punctuation the same from ad to ad. Write numbers and prices the same way every time. ElevenLabs runs a step that turns symbols and numbers into written words, so “5%” and “five percent” can come out differently between tools. Keep a short list of how to spell brand and product names for the voice.

4. The same audio finishing. Finishing is what happens to the file after the voice is made. Set one loudness level, one music level under the voice and one export format.

Test the setup before a campaign. Make the same script three times and listen to the three files in a row. Realistic AI lip sync explains why the audio should be final before the picture is made.

Voice cloning is legal when the person whose voice it is has agreed, and when the result does not deceive people. The rules below come from each law’s own text. This is general information, not legal advice.

New York. A business that uses the “voice of any living person” for advertising “without having first obtained the written consent of such person” is guilty of a misdemeanor (New York Civil Rights Law section 50, October 2026).

California. A person who knowingly uses another’s voice for advertising “without that person’s prior consent” is liable for damages. The amount is the greater of $750 or the actual damages (California Civil Code section 3344, October 2026). The governor approved SB 1111 on 30 September 2026. It adds that “a voice or likeness includes a digital replica” (California SB 1111, October 2026).

Tennessee. The ELVIS Act was signed on 21 March 2024 and adds “voice” to the state’s personal rights law (tn.gov, October 2026). A law firm summary says the act took effect on 1 July 2024 and covers a person’s actual voice and simulations of it (skadden.com, October 2026).

The FTC. In February 2024 the FTC finalized a rule against impersonating governments and businesses. On the same day it proposed to extend the rule to the impersonation of individuals, and its chair named voice cloning as a reason (ftc.gov, October 2026). The FTC’s rule page listed no final rule for individuals when we read it in October 2026 (ftc.gov, October 2026).

The EU. Article 50 of the EU AI Act has applied since 2 August 2026. The European Commission defines a deep fake as AI-made or AI-edited “image, audio or video content that resembles existing persons” and would falsely appear authentic. A business that publishes a deep fake must disclose it when a person first meets it, for example “with visible or audible labels” (European Commission FAQ, October 2026).

A written release makes these rules easier to meet. A useful voice release states five things:

  • Which products and channels the voice may appear in.
  • Which languages the voice may speak.
  • How long the brand may use the clone.
  • Whether the speaker approves each script.
  • What happens to the clone when the contract ends.

What voice vendors ask before a clone

Vendors add their own consent steps on top of the law. Read them before you choose a tool, because they decide whose voice you can clone.

ElevenLabs. For an instant clone, the user confirms that they “have the right and consent to clone the voice”. A professional clone is limited to your own voice: “Even with their consent, you cannot clone someone else’s voice.” The speaker makes and verifies the clone in their own account and can then share it by a private link. The company also says it blocks the cloning of celebrity voices and other high-risk voices (elevenlabs.io/docs and elevenlabs.io/safety, October 2026).

Microsoft. Azure personal voice needs “explicit consent from the user”, given as a recorded statement from the speaker (learn.microsoft.com, October 2026).

For an agency, this means the founder or the actor may have to create the voice in their own account.

Platform disclosure rules for a synthetic voice in ads

Each platform treats an AI voice in its own way. All three rules below were read on the platform’s own page.

PlatformWhen a voice needs a labelWhat the platform does
TikTok adsAudio that is fully AI-made, or voice cloning that makes the main subject say something they did not sayRejects or restricts an ad with undisclosed AI content
Meta adsRealistic-sounding AI audio in ads about social issues, elections or politicsRejects the ad if the disclosure is missing
YouTube uploadsRealistic content that makes a person appear to say something they did not sayNo label needed when you clone your own voice

TikTok. The ads policy asks for the AIGC label or your own clear disclaimer. TikTok also does not allow misuse of the “likeness (audio or visual)” of a public figure without permission (ads.tiktok.com, October 2026).

Meta. Advertisers disclose only in ads about social issues, elections or politics. Meta also says that from 1 June 2026 it uses automated detection on ad media. When it detects AI content, it adds an “AI Info” label in the “About this Ad” menu (transparency.meta.com, October 2026).

YouTube. The help page lists “Cloning one’s own voice to create voice overs or dubs” among uses that need no disclosure. It asks for disclosure when content makes it “appear as if someone gave advice that they did not actually give” (support.google.com, October 2026). This page covers videos that creators upload. Ads on YouTube follow Google Ads policies.

One state law is narrower than it looks. New York’s synthetic performer law asks for a disclosure in ads, and it excludes audio advertisements (New York General Business Law section 396-b, October 2026). The consent duty in section 50 still applies to a cloned voice.

AI disclosure in advertising covers the label rules for images and video. AI lip sync for video ads covers the rules when a real face speaks new words.

Brand voice in a saved workflow

DesignerBox is AI creative production for brands and agencies. It makes images, ads and video from your brand, your rules and your products. Anyone can make an AI picture. Making hundreds that still look like your brand is the hard part.

The same is true for sound. DesignerBox has a text to speech step with 20 voices, a transcription step and a talking avatar with lip sync. You choose the voice once when you build a workflow. Every run after that uses the same voice and the same settings, so ad twenty sounds like ad one.

Cloning a custom voice is outside its scope. If your brand needs a founder’s voice, make the audio in a voice cloning tool with that person’s consent. Then add the audio file to the timeline in the video editor, which has several tracks, transitions, animated text and audio. The full workflow from the first product photo to the finished ad, in one subscription.

The cost is shown before the run. An 8-second clip costs 40 to 560 credits, depending on the model.

Here are the limits. DesignerBox does not post to an ad account. You download the results, or send them with a webhook or an S3 step. Every plan below Ultra is one seat. Uploading your own photos and the commercial license start on the Pro plan. AI video, virtual try-on, upscaling, the image editor and the video editor start on the Premium plan. Plans and credits are on the pricing page.

One voice for every client

An agency builds one workflow per client, with that client’s voice, brand rules and formats saved in it. See DesignerBox for agencies

FAQ

What is AI voice cloning?

AI voice cloning is the use of AI to make a digital copy of one person’s voice. The tool learns the voice from a recording. It then reads any typed script in that voice.

It is legal with the speaker’s consent and without deception. New York requires written consent before a living person’s voice is used in advertising. California requires prior consent. The EU asks for a clear disclosure when audio is a deep fake of a real person (nysenate.gov, leginfo.legislature.ca.gov and digital-strategy.ec.europa.eu, October 2026).

How much audio do I need to clone my voice?

It depends on the tool and the kind of clone. ElevenLabs recommends about 1 to 2 minutes of clear audio for an instant clone and at least 30 minutes for a professional clone (elevenlabs.io/docs, October 2026).

Can I clone another person’s voice?

Only with that person’s consent, and some tools do not allow it at all. ElevenLabs limits a professional clone to your own voice, and the speaker has to verify it (elevenlabs.io/docs, October 2026). For advertising, get the consent in writing.

How do I keep the same AI voice across videos?

Fix four things. Use the same saved voice ID. Reuse the same settings and the same speech model. Write every script in the same style. Finish every audio file at the same loudness and in the same format.

Do I need to label a cloned voice in an ad?

On TikTok, yes, when the audio is fully AI-made or when voice cloning makes the main subject say something they did not say. Meta asks advertisers for disclosure in ads about social issues, elections or politics (ads.tiktok.com and transparency.meta.com, October 2026).

Does DesignerBox clone a custom voice?

No. DesignerBox has text to speech with 20 voices, transcription and a talking avatar with lip sync. To use a cloned voice, make the audio in a voice cloning tool and add the file in the video editor.

Sources

  • ElevenLabs documentation, instant voice cloning, professional voice cloning and text to speech settings (elevenlabs.io/docs), October 2026
  • ElevenLabs safety page: blocked voices and verification (elevenlabs.io/safety), October 2026
  • Microsoft, personal voice overview: sample length and recorded consent statement: learn.microsoft.com, October 2026
  • Higgsfield, guide to a consistent AI voice, published August 2026 (higgsfield.ai), October 2026
  • New York Civil Rights Law section 50: nysenate.gov, October 2026
  • New York General Business Law section 396-b: nysenate.gov, October 2026
  • California Civil Code section 3344: leginfo.legislature.ca.gov, October 2026
  • California SB 1111, Chapter 862, Statutes of 2026: leginfo.legislature.ca.gov, October 2026
  • Tennessee, Governor Lee signs the ELVIS Act (21 March 2024): tn.gov, October 2026
  • Skadden, summary of the ELVIS Act (April 2024): skadden.com, October 2026
  • FTC, proposal on impersonation of individuals (15 February 2024): ftc.gov, October 2026
  • FTC, impersonation rule page: ftc.gov, October 2026
  • European Commission, FAQ on transparency obligations under Article 50 of the AI Act: digital-strategy.ec.europa.eu, October 2026
  • TikTok ads policy, misleading and false content: ads.tiktok.com, October 2026
  • Meta ad standards, ads about social issues, elections or politics: transparency.meta.com, October 2026
  • YouTube Help, disclosing use of AI content: support.google.com, October 2026
  • DesignerBox plans and feature gating: DesignerBox pricing page (designerbox.ai/pricing), October 2026

Laws, platform policies and vendor features checked against each source’s own page as of October 2026. This is general information, not legal advice. Individual results vary.

Vytas

Vytas

Founder at DesignerBox

Vytas is a founder at DesignerBox. He writes about turning creative work a team repeats every week into a system: how a job gets built once, run across a whole catalog, and reviewed in one pass.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Scale your content with AI. Keep your brand.

Build the job once with your brand and your products. Run it on your whole catalog, and see the cost before each run.

One workflow for every product. You see the cost before each run.