Transcribe audio to text
Turn a recording into a clean, editable text transcript — no typing required.
More accurate, slower — recommended
The first run downloads the speech model to your browser (cached after that). Best on desktop — mobile devices may be slower.
Drop a video or audio file to transcribe
or click to browse
How it works
Drop a video or audio file in. WavyVid decodes the audio track right in your browser.
An AI speech-recognition model (Whisper, running fully on-device via WebAssembly/WebGPU) transcribes the audio into timestamped text.
Edit the transcript if needed, then export as .srt/.vtt, or burn the captions directly into your video.
Interviews, meeting recordings, voice memos, and podcast episodes are far easier to search, quote, and repurpose as text. This transcribes the full audio track and gives you a plain-text export alongside timestamped SRT/VTT files, useful for show notes, articles, or accessibility.
Frequently asked questions
Can I export just plain text, without timestamps?
Yes — a plain .txt export is available alongside the timestamped SRT/VTT formats.
How long of a recording can I transcribe?
No hard limit, though very long files take proportionally longer to process and will trigger a heads-up about processing time on large files.
Can I fix mistakes in the transcript?
Yes — every line of the transcript is editable directly in the browser before you export.
Does this work for multi-speaker recordings?
It transcribes all speech into a single stream without speaker labels — useful for content, less so if you specifically need speaker-separated transcripts.