Paste your text, pick a voice, hit generate, download the WAV. The whole thing runs in your browser — no account, no quota, nothing uploaded.

That's the workflow. The difference between robotic output and something you'd actually publish comes down to a few small choices.

Pick the voice first

Before you touch anything else, audition a couple of voices on the same sentence. Voices differ more than the labels suggest — one will sound warm and conversational, another flat and newsreader-ish. The right one for a meditation script is the wrong one for a product demo.

Read one real sentence from your script, not "hello world." You're judging how it handles your punctuation and rhythm, not a generic phrase.

Pace it for the ear, not the eye

Text that reads fine on screen often sounds rushed out loud. Two levers help:

  • Speed. Leave it on Normal for most things. Drop to Slower for tutorials and anything with numbers or names people need to catch; go Faster only for casual, familiar content.
  • Punctuation. Commas and periods become pauses. A wall of text with no commas will sound breathless. Break long sentences, and a line break between paragraphs gives the listener a beat to absorb.

A quick test: close your eyes and listen once. If you lost the thread anywhere, that's where to add a pause.

Export and use it

Generated audio downloads as a WAV file — high quality, but large. If you're posting it somewhere with a size limit (a chat app, a slide deck), run it through the Audio Compression tool afterward to get a smaller MP3 without an audible drop.

Why local matters

Everything happens on your device — the model runs in the browser, so your script never goes to a server. For unreleased content, client copy, or anything you'd rather not paste into a cloud tool, that's the whole point. It also means no per-character billing and no sign-up.


Ready to hear it? Drop your text into the Text to Speech tool and try two voices side by side.