Speech to Text
Cloud transcription · 30+ languages with timestamps
Frequently Asked Questions
Which languages does transcription support?
It handles many languages with automatic language detection, so you don't have to specify the language up front. That covers everything from meeting notes and interviews to lecture recordings in mixed languages.
Can I export subtitles?
Yes. It produces sentence-level timestamps and exports to SRT for subtitles or plain TXT for a clean transcript, both in one click. The timestamps make it easy to jump straight to the moment you need.
Does it support video, and how is it priced?
Yes — upload a video and the audio track is extracted and transcribed automatically. The cloud transcription is billed by duration through redeem codes with no account, so you only pay for the minutes you actually process.
Is there a private mode that doesn't upload my audio?
Yes. There's an opt-in local mode that runs the Whisper model directly in your browser using WebGPU, so the audio never leaves your device. It's the best choice for sensitive recordings; the trade-off is that speed depends on your own hardware.
What audio and video formats can I transcribe?
Common audio formats are supported, and you can also feed in a video file — the audio track is pulled out for you automatically. That makes it easy to caption a clip without converting it first.
Does it work on long recordings?
Yes. The cloud mode comfortably handles long files and is billed by duration, so a two-hour recording is no problem. In local mode, longer files simply take more time depending on your device's speed.
How accurate is the transcription?
It uses modern speech-recognition models with automatic language detection, and accuracy is highest on clear audio with little background noise. Clean recordings from a decent microphone transcribe very reliably; heavy noise, crosstalk or strong accents can lower accuracy.