Audio to Text Transcriber
Turn recordings into text and subtitles with Whisper, privately on your own device.
Recording
Transcription
Choosing it helps with short or mixed recordings.
Model: Whisper tiny (multilingual), running on this device with WebAssembly. It downloads once (about 53 MB) and your recording never leaves your device.
Transcript
About the Audio to Text Transcriber
Choose an audio or video file — a voice note, an interview, a lecture, a meeting or a video — and get a transcript with timestamps, ready to copy or to save as text, SRT or WebVTT subtitles. The speech is recognised by Whisper, OpenAI’s open speech-recognition model, running inside your browser with ONNX Runtime and WebAssembly. Your file is never uploaded: it is decoded, cut into 30-second windows and transcribed on your device.
The page uses Whisper tiny, the smallest multilingual Whisper model (39 million parameters, 99 languages). It detects the spoken language by itself, or you can choose it. It is quick — on a recent laptop, a minute of speech takes roughly 4 to 10 seconds — and good for clear speech in widely spoken languages; it makes more mistakes with noise, strong accents, crosstalk, names and rarer languages, so read the result through before you rely on it.
How to use it
- Choose or drop an audio or video file (MP3, WAV, M4A, AAC, MP4, MOV, WebM, OGG, Opus or FLAC).
- Leave Language on Detect automatically, or choose the language spoken — that is more reliable for short or mixed recordings.
- Select Transcribe. The first time, the speech model (about 53 MB) downloads once; then the transcript appears window by window, with an estimate of the time left. You can stop at any time and keep what is done.
- Select a timestamp to hear that part, then copy the text or download it as Text, Text with times, SRT, WebVTT or JSON.
Examples
jfk.wav — an excerpt of John F. Kennedy’s inaugural address (1961)
1 00:00:00,000 --> 00:00:11,000 And so my fellow Americans ask not what your country can do for you, ask what you can do for your country.
Language detected as English; transcribed on the device in about a second.
Common uses
- Turning interviews, voice notes and lectures into searchable notes.
- Making SRT or WebVTT subtitles for a video, then fixing the few wrong words in a subtitle editor.
- Getting minutes or a summary started from a meeting recording without uploading it anywhere.
- Transcribing recordings that must stay private, such as patient, client or legal conversations.
How the transcription works
- Decoding: WAV, AIFF and FLAC files are decoded by your browser; MP3, M4A/AAC, Opus, OGG and the sound of videos by the open-source Mediabunny library, packet by packet, which needs far less memory for long recordings than decoding the whole file at once (the browser decodes those too if Mediabunny cannot). The sound is mixed down to mono and resampled to 16 kHz with a low-pass filter.
- Recognition: Whisper reads 30-second windows as log-mel spectrograms and writes text with timestamps every 0.02 seconds. The next window starts where the last complete sentence ended, so words are not cut in half. If a window comes out repetitive or unlikely, it is decoded again with a little randomness (as the reference implementation does), and windows the model rates as silence are skipped.
- Language: the first window decides the language when you leave it on automatic.
- Privacy: the model files come from this site and are cached by your browser; the audio and the transcript stay in the page. Closing the tab clears them.
Accuracy and speed
Whisper was trained on 680,000 hours of speech (Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision). The tiny model trades accuracy for speed and size: it is best with one clear speaker and little background noise, in languages with a lot of training data such as English, Spanish, German, French or Japanese. With Hindi speech it may write an English translation, or the words in Latin letters, instead of Hindi; the page warns you whenever a transcript is not in the script of the language spoken. For better results, record close to the microphone, choose the language instead of detecting it, and trim long silences.
Transcription runs on your processor with one WebAssembly thread: on a recent laptop often more than 10 times faster than real time, on phones slower. A one-hour recording usually takes several minutes; keep the tab open while it works.
Limitations
- Speakers are not told apart (no “Speaker 1 / Speaker 2” labels), and timestamps are per sentence, not per word.
- Translation into English is not offered: the small model that can run in a browser translates poorly.
- Recordings up to 2 hours can be transcribed at a time; split longer ones. Very long files need a lot of memory on phones.
- The tiny model can mishear names, numbers and technical terms, and sometimes repeats or invents a phrase in noisy or silent parts. Always proofread.
- In languages with less training data it is much weaker: Hindi speech, for example, may come out as an English translation or in Latin letters. The page warns you when a transcript is not in the script of the language spoken.
- Which formats open depends on your browser’s codecs: a file that will not open can be converted first with the Audio Converter.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. The first time you transcribe, the Whisper speech model and its engine (about 53 MB) are downloaded from this site; your browser keeps them for next time. Your recordings are never uploaded.
Frequently asked questions
Is my recording uploaded to a server?
No. The file is decoded and transcribed in your browser, on your device. Only the speech model itself is downloaded, from this site, the first time you use the tool.
How accurate is it?
For clear speech by one speaker in a widely spoken language, most words come out right, with punctuation and capital letters. Noise, music, accents, several people talking at once, names and rare words cause mistakes. It uses Whisper tiny, the smallest Whisper model, so expect to correct a few words.
Which languages does it understand?
It recognises the 99 languages Whisper was trained on, including English, Hindi, Bengali, Tamil, Telugu, Marathi, Urdu, Gujarati, Kannada, Malayalam, Punjabi, Spanish, French, German, Arabic, Chinese and Japanese, but the tiny model is much more accurate in languages with a lot of training data. Hindi speech, for example, can come out as an English translation; when a transcript is not in the script of its language, the page tells you so you can check it.
Can I make subtitles for a video?
Yes. Open the video file itself (MP4, MOV or WebM); the sound is transcribed and you can download SRT or WebVTT subtitles with timings, which video players, YouTube and editing apps can load.
Why does it download 53 MB first?
That is the Whisper model and the ONNX Runtime engine that runs it. They are downloaded once from this site and kept in your browser’s cache, so later transcriptions start at once — also with a slow connection.
Does it work without an internet connection?
After the model has been downloaded once, yes, as long as your browser keeps it in its cache. The first time needs a connection.