Transcribe audio to text
Turn a recording into a clean, editable text transcript — no typing required.
More accurate, slower — recommended
The first run downloads the speech model to your browser (cached after that). Best on desktop — mobile devices may be slower.
Drop a video or audio file to transcribe
or click to browse
How it works
Drop a video or audio file in. WavyVid decodes the audio track right in your browser, no upload involved.
An AI speech-recognition model (Whisper, running fully on-device via WebAssembly/WebGPU) transcribes the audio into timestamped text, segment by segment.
Edit the transcript if any words came out wrong, then export as .srt/.vtt, or burn the captions directly into the video frame.
When you'd use this
- You recorded an interview or meeting and need a written record to quote from or reference later.
- You want to turn a podcast episode into show notes or a blog post without retyping it by hand.
- You need a searchable text version of a voice memo or lecture recording.
Frequently asked questions
Can I export just plain text, without timestamps?
Yes — a plain .txt export is available alongside the timestamped SRT/VTT formats.
How long of a recording can I transcribe?
No hard limit, though very long files take proportionally longer to process and will trigger a heads-up about processing time on large files.
Can I fix mistakes in the transcript?
Yes — every line of the transcript is editable directly in the browser before you export.
Does this work for multi-speaker recordings?
It transcribes all speech into a single stream without speaker labels — useful for content, less so if you specifically need speaker-separated transcripts.
Is the transcript accurate enough to publish directly?
It's a strong starting point, but reviewing for misheard names, jargon, and homophones before publishing is still worth the few extra minutes — the built-in editor makes that quick.
Want more detail? Read: How to transcribe audio to text for free