Speech to Text
Turn audio or video into a transcript or timed subtitles in your browser. Positioned as Audio/Video → Text/Subtitles — a WASM speech model runs on your device, then you download TXT, SRT, or VTT.
Drag & drop audio or video, or choose a source
MP3, WAV, M4A, MP4, MOV, WEBM — transcript stays in your browser
Audio / video → text & subtitles · private WASM model
Download as
TXT for a plain transcript; SRT / VTT for timed subtitles.
How to use this tool
Drop audio or video
MP3, WAV, M4A, MP4, MOV, and WEBM are supported. The file is read locally in this tab.
Transcribe on-device
A speech recognition model runs in WebAssembly. The first run downloads ~40 MB and caches it; later runs reuse the cache.
Download TXT, SRT, or VTT
Plain transcript or timed cues for editors and players — nothing is uploaded to a server.
Audio/Video → Text/Subtitles
Use Speech to Text when you need a transcript from a podcast, call recording, or camera clip. Prefer SRT/VTT when you will burn captions or sync with a player; use TXT when you only need searchable dialogue.
Private browser processing
Recognition uses on-device Whisper tiny (multilingual). Choose a listed language or Auto-detect — the tiny model is not claimed as high-accuracy for every locale. Pair it with the subtitle editor if cue timing needs a later edit.
Frequently Asked Questions
More on quality loss: conversion quality guides · compression guides
Is this an AI chat feature?
No. Format Convertly treats speech-to-text as a conversion: Audio/Video → Text/Subtitles, the same privacy model as OCR and subtitle converters.
Which outputs can I download?
TXT for a plain transcript, plus SRT and VTT for timed subtitles.
Are files uploaded?
No. Decoding and recognition stay in your browser. Only the public model/runtime assets are fetched on first use.
Need a format-specific landing?
Open MP3 to Text, WAV to Text, M4A to Text, MP4 to Text, MOV to Text, or WEBM to Text under converters.
Ready to convert more formats?
Browse converters or jump back to the tool above.