Transcribe audio to text

Turn a recording into a clean, editable text transcript — no typing required.

More accurate, slower — recommended

The first run downloads the speech model to your browser (cached after that). Best on desktop — mobile devices may be slower.

Drop a video or audio file to transcribe

or click to browse

Processed on your device — never on a server.

How it works

1

Drop a video or audio file in. WavyVid decodes the audio track right in your browser, no upload involved.

2

An AI speech-recognition model (Whisper, running fully on-device via WebAssembly/WebGPU) transcribes the audio into timestamped text, segment by segment.

3

Edit the transcript if any words came out wrong, then export as .srt/.vtt, or burn the captions directly into the video frame.

When you'd use this

  • You recorded an interview or meeting and need a written record to quote from or reference later.
  • You want to turn a podcast episode into show notes or a blog post without retyping it by hand.
  • You need a searchable text version of a voice memo or lecture recording.

Frequently asked questions

Can I export just plain text, without timestamps?

Yes — a plain .txt export is available alongside the timestamped SRT/VTT formats.

How long of a recording can I transcribe?

No hard limit, though very long files take proportionally longer to process and will trigger a heads-up about processing time on large files.

Can I fix mistakes in the transcript?

Yes — every line of the transcript is editable directly in the browser before you export.

Does this work for multi-speaker recordings?

It transcribes all speech into a single stream without speaker labels — useful for content, less so if you specifically need speaker-separated transcripts.

Is the transcript accurate enough to publish directly?

It's a strong starting point, but reviewing for misheard names, jargon, and homophones before publishing is still worth the few extra minutes — the built-in editor makes that quick.

Related tools

Ad slot