How to Clone a Voice from a Short Sample (And Get It Right)
Give it 10–20 seconds of clean speech and it will read new text back in that voice. The quality of the clone is decided almost entirely by that one sample — get the sample right and everything else follows.
What a good sample looks like
The tool wants WAV, MP3 or M4A — mono, at least 24kHz, 10–20 seconds of clear speech (hard caps: 60 seconds, 10 MB). Within that, aim for:
- One speaker, talking normally. Plain narration, not shouting or whispering. The clone copies the energy it hears.
- Dead-quiet background. No music, no TV, no street noise. Whatever's behind the voice gets baked in.
- A few full sentences, not isolated words. The model needs natural rhythm and intonation to imitate.
The mistakes that ruin a clone
Most disappointing results trace back to the sample, not the tool:
- Background noise or echo. A reverby room or a fan hum makes the clone sound muddy. Record somewhere soft and small (a closet full of clothes is a classic trick).
- Clipping. If the original was recorded too loud and the peaks are distorted, the clone inherits the distortion. Back off the mic or input level.
- Two voices. A clip where someone else talks over the end confuses the model. Trim to just your target speaker.
- Too short. Three seconds isn't enough rhythm to copy. Give it the full 10–20.
If a clone sounds off, don't tweak settings — re-record a cleaner sample. That fixes it nine times out of ten.
Generate, then sanity-check
Once it's cloned, type a sentence the original sample didn't contain and listen. New words are the real test — anything can echo back a phrase it was given. The output downloads as a WAV.
Use it responsibly
Only clone a voice you have the right to use — your own, or someone who has clearly agreed. Cloning a person's voice without consent is the kind of thing that gets people into real trouble, and it's not what this is for.
On privacy: the processing runs in your browser, so the sample stays on your device rather than going to a server.
Have a clean clip ready? Drop it into the Voice Cloning tool and have it read something new.