AI voice cloning from a 30-second sample.
Upload or record a 30-second audio sample, then use AI voice cloning to create narration in your own voice with text to speech.
A sample in, a voice out.
Upload or record a clean sample, confirm permission, and create a clone. It appears in the voice picker after processing and activation complete.
Three steps, once.
Available while your account and provider support it.
Read anything at a natural pace in a quiet room. One to two minutes gives a noticeably better result than the thirty-second minimum.
Upload or record a clean sample, confirm permission, and create a clone. It appears in the voice picker after processing and activation complete.
Your clone sits in the same picker as the 440+ catalog voices, with the same speed, pitch, and emotion controls.
Yours, private, and reversible.
Thirty seconds is the floor, not the target. Most people get what they want from about ninety seconds of ordinary reading.
Clones are scoped to your account. They never appear in the public catalog and no other user can generate with them.
Choose MP3, WAV, or FLAC in the studio. Available audio settings depend on the selected model.
The clone keeps how you actually speak rather than flattening you to a neutral accent. Pick the language you recorded in.
Three clones on Starter, ten on Creator, thirty on Pro. Enough for a personal voice, a brand voice, and a few experiments.
Deleting a clone removes it from your account and frees its slot. Sample and preview cleanup may be queued. Provider-side voice-model deletion cannot currently be guaranteed.
Your voice stays yours.
Cloning is the one feature where the rules matter more than the specs. Ours are short and we enforce them.
Clone only voices you own or have explicit permission to use. You confirm this before the clone is created, and it is a condition of the account.
A clone is visible only inside your account. It is not published, not shared between users, and not added to the catalog.
Deleting a clone removes it from your account and frees its slot. Sample and preview cleanup may be queued. Provider-side voice-model deletion cannot currently be guaranteed.
Who clones a voice, and why.
Publishing three shorts a week means three recording sessions, three level matches, and three chances to sound tired. A clone removes the session entirely — change one line in the script and regenerate that line alone, still in your voice, still matching the rest.
Narrating your own book is the version readers want, but a full manuscript is days of studio time and a voice that tires by chapter four. With a clone you narrate a chapter at a time, at whatever hour suits, and chapter twelve sounds exactly like chapter one.
One brand voice across product videos, help centre clips, and phone menus — without rebooking the same reader every quarter. Creator gives you ten slots, so a brand voice, a support voice, and a few regional variants can coexist.
What cloning is, and what makes it good.
How voice cloning works
A clone is built from a short recording of you speaking. The system learns the characteristics that make your voice recognisable — timbre, resonance, how you shape vowels, your habitual pace — and produces a voice model from them.
From then on you type text and the model speaks it. It is not replaying your recording; it is generating new speech in your voice, so you can say sentences you never recorded. The sample teaches the model what you sound like, not what you said.
What makes a good sample
A quiet room matters more than an expensive microphone. Background hum, a fan, or traffic gets learned as part of your voice, and it will follow you into every generation afterwards.
Read at your normal pace and let your voice move — a flat, careful read produces a flat clone. Thirty seconds is the minimum; one to two minutes of varied speech is where quality really lands. If a clone sounds lifeless, the sample is almost always the reason.
Clone, re-record, or hire.
A clone will not out-act a professional reader on a difficult script. It wins on the boring dimensions, which is usually what decides the work.
| Clone your voice | Re-record each time | Hire a reader | |
|---|---|---|---|
| Setup time | 30 seconds is the minimum sample length, not a processing-time guarantee. | Mic and room setup every session | Casting, briefing, and a contract |
| Cost per revision | Nothing beyond the characters used | Your time, plus matching old levels | Often a new session fee |
| Consistency | Use the same voice; delivery can vary. | Changes with your cold, room, and mood | Consistent while the reader is available |
| Availability | Available while your account and provider support it. | Whenever you can get to a quiet room | Subject to their calendar |
The short version.
Pick a plan when you're ready.
Cloning needs a paid plan — the free tier does not include clone slots. The number that matters here is how many voices you can keep.
Every account starts with a one-time allowance of 500 free characters that never refills — no card required.For hobbyists getting started with AI voice.
/ month
For creators publishing every week.
/ month
For teams and heavy production.
/ month
- Studio HD and Turbo voice models
- MP3, WAV and FLAC downloads
- Emotion, speed and pitch control
- Sound cues and pronunciation dictionary
Questions about cloning.
Thirty seconds is the minimum the system accepts. One to two minutes gives a noticeably better clone, because the extra material is what teaches it your pacing and range rather than just your timbre. Beyond about three minutes the gains flatten out.
Yes — record your sample in the language you want the clone to speak, and it keeps your natural accent rather than flattening it. The catalog spans 40 languages, and clones follow the same engine.
Only you. A clone is scoped to your account: it never appears in the public catalog, other users cannot see or select it, and it is not shared between accounts.
Deleting a clone removes it from your account and frees its slot. Sample and preview cleanup may be queued. Provider-side voice-model deletion cannot currently be guaranteed.
Almost always the sample. A careful, monotone read in a room with a fan produces a careful, monotone clone with a hum baked in. Re-record ninety seconds of natural, varied speech in a quiet space and build a fresh clone — that fixes it in most cases.
Three slots on Starter, ten on Creator, and thirty on Pro; Lifetime includes fifty. The free tier does not include clone slots. Deleting a clone frees its slot instantly.
Samples are sent to the voice provider to create your clone. Provider retention and any training use depend on its terms and configuration; provider-side erasure is not guaranteed.