ElevenLabs alternatives fall into four jobs: voiceover for ads and video, voice cloning with consent, real-time voice for apps, and open-source text to speech you host yourself. A fifth route is a video model that returns audio with the picture. Pick the job first. Then check languages, cloning consent, commercial-use terms, API access and how the audio reaches your timeline.
A list of ElevenLabs competitors in one column mixes a narration editor, a low-latency API for phone bots and a model you run on your own GPU. Those three do not replace each other.
This guide is for an agency or a brand team that needs voiceover for video ads and product videos. Every fact comes from the vendor’s own page or model card, read on 4 October 2026. If you only need to choose how an ad gets its voice, start with the three voiceover paths for UGC ads.
Key Takeaways
- Four jobs, one name. An ElevenLabs alternative can mean ad voiceover, voice cloning, real-time voice or a self-hosted model. Each job has a different first check.
- ElevenLabs covers a lot. Its docs list Eleven v4 with 90+ languages and Flash v2.5 at about 75ms. A switch should solve a named problem.
- Cloning starts with consent. OpenAI, Microsoft and ElevenLabs all ask for a recorded statement or a verification step before a custom voice exists.
- Open source has license traps. Kokoro’s weights are Apache-licensed. F5-TTS’s pre-trained models are CC-BY-NC, which rules out a paid ad.
- Video models can replace the voice file. Veo 3.1 returns 8-second clips with natively generated audio, so some ads need no voice tool at all.
What ElevenLabs does well
ElevenLabs is a wide audio platform. Its text to speech page lists over 11,000 voices in its Voice Library, plus Voice Design and Voice Cloning (elevenlabs.io/text-to-speech, October 2026). Its product menu also lists a voice changer, sound effects, a voice isolator, dubbing, music and a platform for voice agents.
The model docs list four generations side by side (elevenlabs.io/docs/overview/models, October 2026):
- Eleven v4: 90+ languages and audio tags for fine control.
- Eleven v3: 70+ languages, built for expressive delivery and multi-speaker dialogue.
- Multilingual v2: 29 languages, described as the most stable on long-form reads.
- Flash v2.5: 32 languages at about 75ms. ElevenLabs notes that this number is model inference time only.
It offers two kinds of cloning. Instant Voice Cloning works from 1 to 5 minutes of audio. Professional Voice Cloning asks for 30 minutes or more, and includes a voice verification step (elevenlabs.io/voice-cloning, October 2026).
There are four common reasons to look elsewhere: a different job, a rule about where data may go, a need to own the model, or a workflow where the voice should arrive inside the video.
Four jobs, and what to check first for each
The word “alternative” hides the job. Name the job, and the list gets short.
| Job | Who has it | Check first | Options covered here |
|---|---|---|---|
| Ad and video voiceover | Agencies, brand teams | Commercial-use terms | Murf, WellSaid, OpenAI text to speech, Gemini TTS |
| Voice cloning with consent | Brands with a founder or spokesperson voice | Consent and rights | Fish Audio, Azure personal voice, OpenAI custom voices |
| Real-time voice | Developers of voice agents and apps | Latency and languages | Cartesia, Hume |
| Open-source models you host | Teams with an engineer and a GPU | The model license | Kokoro, Chatterbox, XTTS-v2, F5-TTS |
| Native audio from a video model | Ad teams making short clips | Clip length | Veo 3.1 |
Sources: each vendor’s own page or model card, listed under Sources, read on 4 October 2026.
An ad team mostly works in rows one and five. Rows two to four matter when the brief changes.
Job 1: voiceover for ads and video
You paste a script, pick a preset voice, and export an audio file for the edit.
Murf. Murf’s text to speech page lists 200+ voices and states “full commercial usage rights for audio created with our text-to-speech converter” (murf.ai/text-to-speech, October 2026). Its homepage lists dubbing and a voice-over-video feature next to the editor, and its Falcon 2 model covers 35+ languages (murf.ai, October 2026).
WellSaid. WellSaid lists 280+ voices across 18 languages with regional variants. Its site says every voice in the library includes full commercial usage rights, and it names paid advertising as a use (wellsaid.io, October 2026). The language list is shorter than the one ElevenLabs publishes, so check it against your markets.
OpenAI text to speech. OpenAI’s Audio API has a speech endpoint with 11 built-in voices. Its guide adds one rule an ad team must plan for: “Our usage policies require you to provide a clear disclosure to end users that the TTS voice they are hearing is AI-generated and not a human voice” (OpenAI text to speech guide, October 2026). It is an API, so someone has to write the call.
Gemini TTS. Google’s speech generation docs say Gemini 3.8 Flash TTS supports over 130 languages and detects the input language automatically. The page lists 30 featured voices (Gemini API speech generation, October 2026). This is also an API.
If the ad needs music or a sung line, see our guide to the advertising jingle.
Job 2: voice cloning with consent
Cloning fits when the voice is a brand asset, such as a founder who already fronts the account. Check consent before any feature.
Fish Audio. Fish Audio’s homepage says it clones a voice from a 15-second clip, speaks 30+ languages with any voice, and hosts over 2,000,000 voices, many uploaded by users. Its FAQ states that the free plan is for personal use only and that commercial use needs a paid plan (fish.audio, October 2026). With a user-uploaded voice, ask who gave the rights to it before it goes into an ad.
Azure personal voice. Microsoft’s docs say a personal voice is created from “a verbal statement and a short speech sample”, and can then speak more than 90 languages. Access to the API is restricted to eligible customers and approved use cases (learn.microsoft.com, Azure Speech personal voice overview, October 2026).
OpenAI custom voices. OpenAI asks for two recordings: a consent recording, in which the voice actor reads a consent phrase, and a sample recording. Custom voices are limited to eligible customers (OpenAI custom voices guide, October 2026).
ElevenLabs itself. Its use policy bans replicating another person’s voice “without consent or legal right” (elevenlabs.io/use-policy, October 2026).
Three legal sources, each checked on 4 October 2026:
- FTC. The FTC’s rule on impersonation of government and businesses has been in effect since 1 April 2024. The agency held an informal hearing on a proposed amendment that would cover impersonation of individuals (FTC rule page, October 2026).
- Tennessee. Governor Bill Lee signed the ELVIS Act on 21 March 2024. The state says it updates Tennessee’s personal rights law to protect a performer’s voice from misuse of AI (Tennessee Governor’s office, October 2026).
- EU. Article 50 of the AI Act has applied since 2 August 2026. The Commission’s FAQ defines a deepfake as AI-generated image, audio or video content that resembles existing persons and would falsely appear authentic, and such content must be disclosed (European Commission FAQ, October 2026).
This is general information, not legal advice. For a face that speaks, the consent questions are wider, and our guide to lip sync in video ads covers them.
Jobs 3 and 4: real-time voice and open-source models
Real-time voice for agents and apps
This job belongs to developers. A voice agent on a phone line needs the first sound in a fraction of a second. An ad team exporting a 20-second read does not.
Cartesia. Cartesia’s Sonic page describes a real-time text to speech API. It says you can clone a voice with 10 seconds of audio and localize it into 44 languages (cartesia.ai/sonic, October 2026). Its docs add a Pro Voice Clone trained on 30 minutes or more of audio (docs.cartesia.ai, October 2026).
Hume. Hume’s docs describe Octave as a text to speech system built on a language model, so it reads the meaning of the text. Octave 2, in preview, lists 11 languages and model latency of about 100ms, and clones a voice from as little as 15 seconds of audio. The docs state that you retain full ownership of the audio you generate (dev.hume.ai, October 2026).
ElevenLabs Flash v2.5 and Murf Falcon 2 sit in this group too. Latency does not change how a finished ad sounds, so an ad team can skip this row.
Open-source text to speech you host
An open-source text to speech model runs on your own hardware. Nothing leaves your network, and there is no per-character meter. You need an engineer, and you must read the license, because “open” does not always mean “free for ads”.
| Model | Published by | License as written on its page | What the page lists |
|---|---|---|---|
| Kokoro-82M | hexgrad, on Hugging Face | Apache 2.0 | 82 million parameters. Version 1.0 lists 8 languages and 54 voices |
| Chatterbox | Resemble AI, on GitHub | MIT | Voice cloning, a multilingual model with 23 languages, a watermark in every file |
| XTTS-v2 | Coqui, on Hugging Face | Coqui Public Model License | Cloning from a 6-second clip, 17 languages |
| F5-TTS | SWivid, on GitHub | Code MIT, pre-trained models CC-BY-NC | Multi-speaker generation |
Sources: Kokoro-82M model card, Chatterbox repository, XTTS-v2 model card and F5-TTS repository, read on 4 October 2026.
The F5-TTS repository says its pre-trained models are CC-BY-NC “due to the training data”. NC means non-commercial, so those weights do not fit an ad. XTTS-v2 uses its own license, so read the Coqui Public Model License before any client work.
The fifth route: audio that arrives with the video
Some video models return speech and sound with the picture, so no voice file exists and nothing needs syncing. The guide to AI video generators with sound lists them.
Google’s docs describe Veo 3.1 as “a model for generating 8-second videos (720p, 1080p, or 4k) with natively generated audio” (Gemini API Veo guide, October 2026). Plan around the date as well. Google lists 22 October 2026 as the earliest shutdown date for the Veo 3.1 preview models in the Gemini API, and names Gemini Omni Flash as the replacement (Gemini API deprecations, October 2026).
Eight seconds holds a hook or one product line. A 45-second read still needs a voice file.
The other built-in route is a presenter. An avatar platform renders the voice and the face together, and our comparison of HeyGen alternatives sorts those tools.
Five checks before you switch
Run these five on the vendor’s own pages, with one real script from a live campaign.
- Language coverage. List the markets you run ads in this quarter. Check that each language is on the model you would use, since one vendor’s models differ. ElevenLabs lists 29 languages on one model and 90+ on another.
- Cloning consent and rights. Ask who recorded the voice, and whether you hold their consent in writing. A preset voice from the vendor’s own library avoids this step.
- Commercial-use terms. Find the sentence that covers paid ads, on the plan you would buy. Fish Audio limits its free plan to personal use. Check any disclosure duty too, such as the one in OpenAI’s guide.
- API or batch access. One read needs an editor. Forty product videos in six languages need an API or a batch route, and someone to run it.
- How the audio reaches the timeline. Count the steps from script to finished cut: export, upload, place, trim, caption. Every manual step repeats for every variant. Our guide to editing AI-generated video covers the cut itself.
Voice inside a DesignerBox workflow
DesignerBox is AI creative production for brands and agencies. It is not a voice tool, and it does not clone voices. If you need a cloned founder voice or a real-time voice API, one of the options above fits better.
What it has is the voice as a step. A text to speech step offers 20 voices. A transcription step turns speech into text for captions. A talking avatar with lip sync speaks a script. Some of its video models return audio with the picture, which is the fifth route above. The video editor is a real timeline, with several tracks, transitions, animated text and audio, so the read, the clip and the captions meet in one place.
Anyone can make an AI picture. Making hundreds that still look like your brand is the hard part. The same holds for sound. You set the voice, the brand rules and the framing once in a workflow, and the workflow reads them on every run. Ad forty then sounds like ad one. The AI video ads page shows the route from product photo to finished clip.
The full workflow from the first product photo to the finished ad, in one subscription. The cost is shown before the run.
The limits, stated plainly:
- 20 preset voices, no cloning. A dedicated voice platform lists far more voices and languages.
- No public API. DesignerBox has 68 tools over MCP, so an AI chat such as Claude, ChatGPT or Cursor can run your workflows.
- No publishing to a channel. You download the results, or send them with a webhook or an S3 step.
- One seat below Ultra. Team features, shared brand kits and white label are on the Ultra plan.
Uploading your own photos and the commercial license start on the Pro plan. AI video, virtual try-on, upscaling, the image editor and the video editor start on the Premium plan. Plans and credits are on the pricing page.
There is a free plan, and it runs on sample products. Start from a template and run it once before you move a whole campaign.
FAQ
What is the best ElevenLabs alternative for video ads?
It depends on the job. For a preset voice read with stated commercial rights, Murf and WellSaid both publish that term on their own sites (murf.ai and wellsaid.io, October 2026). For a short clip where the voice arrives with the picture, a video model with native audio such as Veo 3.1 removes the voice file. For a presenter, use an avatar platform.
Is there an open-source alternative to ElevenLabs?
Yes, several. Kokoro-82M has Apache-licensed weights, and Chatterbox from Resemble AI uses the MIT license. XTTS-v2 uses the Coqui Public Model License. F5-TTS has MIT code, and its pre-trained models are CC-BY-NC, which is non-commercial. Read the license on the model’s own page before you use any of them in a paid ad.
Can a video model replace a voice tool for an ad?
For a short clip, yes. Google describes Veo 3.1 as a model for 8-second videos with natively generated audio, so the voice arrives with the picture. Eight seconds holds a hook or one product line. A longer read still needs a voice file placed on a timeline.
Do I need consent to clone a voice for an ad?
Yes. ElevenLabs bans replicating another person’s voice without consent or legal right. OpenAI asks for a consent recording from the voice actor, and Microsoft asks for a verbal statement. State and EU rules add duties on top, so check every market the ad runs in.
Does DesignerBox replace ElevenLabs?
No. DesignerBox does not clone voices and is not a standalone voice tool. It has a text to speech step with 20 voices, transcription, a talking avatar with lip sync and a video editor with audio tracks. In DesignerBox the cost is shown before the run. AI video and the video editor start on the Premium plan.
Sources
All read on 4 October 2026.
- ElevenLabs voice library, Voice Design and the product list: (elevenlabs.io/text-to-speech, October 2026)
- ElevenLabs models, language counts and Flash v2.5 latency: (elevenlabs.io/docs/overview/models, October 2026)
- ElevenLabs Instant and Professional Voice Cloning: (elevenlabs.io/voice-cloning, October 2026)
- ElevenLabs rule on replicating a voice without consent: (elevenlabs.io/use-policy, October 2026)
- Murf voices, commercial usage rights, Falcon 2 languages, dubbing and voice over video: (murf.ai and murf.ai/text-to-speech, October 2026)
- WellSaid voices, languages, commercial usage rights and paid advertising use: (wellsaid.io, October 2026)
- Fish Audio cloning, languages, voice count and free-plan terms: (fish.audio, October 2026)
- Cartesia Sonic cloning and languages: (cartesia.ai/sonic, October 2026). Pro Voice Clone: (docs.cartesia.ai, October 2026)
- Hume Octave languages, latency, cloning and ownership: (dev.hume.ai/docs/text-to-speech-tts/overview, October 2026)
- Azure personal voice, verbal statement, languages and restricted access: (learn.microsoft.com/en-us/azure/ai-services/speech-service/personal-voice-overview, October 2026)
- OpenAI built-in voices and the disclosure rule: OpenAI text to speech guide
- OpenAI consent recording and eligibility: OpenAI custom voices guide
- Gemini 3.8 Flash TTS languages and voices: Gemini API speech generation
- Veo 3.1 clip length and native audio: Gemini API Veo guide
- Veo 3.1 preview shutdown date and replacement: Gemini API deprecations
- Kokoro license, size, languages and voices: Kokoro-82M model card
- Chatterbox license, languages and watermark: Chatterbox repository
- XTTS-v2 license, languages and 6-second cloning: XTTS-v2 model card
- F5-TTS code and model licenses: F5-TTS repository
- FTC impersonation rule and the proposed amendment on individuals: FTC rule page
- Tennessee ELVIS Act signing: Tennessee Governor’s office
- EU AI Act Article 50, application date and deepfake definition: European Commission FAQ
- DesignerBox video editor, text to speech, transcription and plan gates: DesignerBox video editor page and pricing page (designerbox.ai), October 2026
Tool features checked on each vendor’s own page on 4 October 2026. This article prints no competitor prices; each vendor lists its plans on its own pricing page. Model licenses are quoted as written on each model’s page and can change. DesignerBox publishes this article and is not a voice cloning tool. This is general information, not legal advice.