Audio to Text (Transcription)
Turn voice notes, meetings and lectures into text, on your own device.
Free preview.
- Free preview: the transcript of the first half of the recording (up to 60 seconds), the first half of its words (up to 100) with times, and the length of the whole recording.
- Locked until you unlock it: download and copy.
- Unlock: Pro pass, ₹179 for 30 days, a one-time payment that never renews.
Ways to unlock shows how to get the full result.
Printing this result is locked in the free preview.
Recording
Speech
The model is told this language; choose the one spoken in the recording.
Translates best with Standard or Most accurate.
Choose how to run this
Stopping may still count toward today's free server AI.
Free server AI today (estimate):
Transcript
Locked in the free preview. Opens the ways to unlock this result.
Locked in the free preview. Batch runs unlock with a pass.
Locked in the free preview. Query results unlock with a pass.
About the Audio to Text (Transcription)
Choose a recording — a voice note, a meeting, a lecture, an interview, a podcast or the sound of a video — and get its words as text, split into paragraphs at the pauses, with a time on every paragraph that plays that moment when you select it. Speech recognition runs in your browser on your own device, with Whisper for 99 languages or Moonshine for fast English; the recording is never uploaded, and the browser’s own speech recognition (which sends sound to its maker) is never used.
With a pass you save the transcript as plain text, text with times, a Word document, Markdown, SRT or WebVTT subtitles, or JSON with every segment’s times, and copy it. Without one you get the free preview: the first part of the recording, transcribed on your device.
How to use it
- Choose or drop a recording (MP3, M4A, WAV, OGG, Opus, FLAC, or a video such as MP4, MOV or WebM), up to 2 hours long.
- Choose the language spoken (it starts as your browser’s language), and tick Translate into English to get English text of speech in another language (with Standard or Most accurate: they translate far better than Fast).
- Pick a quality: Fast English (Moonshine tiny), Fast (Whisper tiny), Standard (Whisper base) or Most accurate (Whisper small). The first time, agree to the one-time download; its exact size is shown first.
- Press Transcribe and keep the tab open. Paragraphs appear as they are recognised; select a time to hear that part.
- With a pass, choose a format and press Download or Copy; without one, the free preview shows the first part and Ways to unlock shows how to get the whole transcript.
Examples
jfk.mp4 — the end of John F. Kennedy’s inaugural address
[00:00:00] And so my fellow Americans ask not what your country can do for you, ask what you can do for your country.
Whisper tiny on the device. Small models sometimes leave out commas (here after “so” and “Americans”), so read the text through.
A long recording with pauses between topics
[00:00:00] The first paragraph, up to the first long pause… [00:06:12] The next paragraph…
Each paragraph starts with the time it begins at. A new paragraph starts after a pause of two seconds or more, or at a sentence end once a paragraph is long.
Common uses
- Turning interviews, voice notes and lectures into searchable, quotable text.
- Notes and minutes from meeting recordings that must not be uploaded anywhere.
- A Word document of a recorded talk to edit and share.
- English text from a recording in another language.
- The first draft of subtitles for a video, saved as SRT or WebVTT.
How the transcription works
- Sound: the recording is decoded on your device and mixed to 16 kHz mono; long silences are skipped, and speech is cut into parts of up to 30 seconds at pauses.
- Whisper (open source; tiny 39 million parameters, base 74 million, small 244 million) was trained on 680,000 hours of speech in many languages (Radford et al., “Robust Speech Recognition via Large-Scale Weak Supervision”). It writes punctuation and a time for each sentence or phrase, and can translate into English.
- Moonshine tiny (MIT licence, 27 million parameters, English only) needs about five times less computing than Whisper tiny.en for a 10-second segment with no more errors (Jeffries et al., “Moonshine: Speech Recognition for Live Transcription and Voice Commands”). It gives one time per stretch of speech, not per sentence.
- Both run on your processor with WebAssembly in a background worker, after a one-time download from MySmartCoPilot that your browser keeps. Paragraphs start after pauses of two seconds or more.
Which quality to choose
- Fast English (Moonshine tiny, about 32 MB): English only; quickest.
- Fast (Whisper tiny, about 44 MB): any of the 99 languages; quick, fine for clear speech in widely spoken languages.
- Standard (Whisper base, about 80 MB): better words and punctuation; the default.
- Most accurate (Whisper small, about 252 MB): better with noise, accents, names and languages with less training speech, such as Hindi, Tamil or Bengali; slower, so download it on Wi-Fi.
While it works the page shows how much is done and, after the first part, about how long is left.
The formats
- Text and Text with times (
[00:12:31]before each paragraph). - Word (.docx): a title, the language and length, and the paragraphs with their times, ready to edit in Word, Google Docs or LibreOffice.
- Markdown with the times in bold, for notes apps.
- SRT and WebVTT subtitles, cut to Netflix’s limits (42 characters a line in English and most languages, 16 in Chinese and Korean, 13 in Japanese, 35 in Thai; 2 lines; up to 7 seconds).
- JSON with every segment’s start and end in seconds, for your own programs.
Without a pass, these exports and copying are locked; the free preview shows the first part on the page.
Limitations
- Choose the language spoken: the model is told the language and does not detect it here. With the wrong language the text comes out wrong or translated.
- Speakers are not told apart (no “Speaker 1 / Speaker 2”), and times are per sentence or per stretch of speech, not per word.
- Mistakes happen with names, numbers, technical terms, strong accents, background music and people talking over each other: read the transcript through before you rely on it.
- Up to 2 hours and 2 GB per file; cut longer recordings into parts. AMR voice recordings must be converted first (for example with the Audio Converter).
- Translation is into English only (Whisper’s own translation) and is less accurate than transcription; the Fast model often translates badly.
- Phones are much slower than computers, and the Most accurate model needs a lot of memory.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. The first time, the speech model you choose (about 32, 44, 80 or 252 MB) and its engine are downloaded from MySmartCoPilot after you agree, and your browser keeps them. Your recording never leaves your device.
Frequently asked questions
What do I get without a pass?
Without a pass, Audio to Text (Transcription) shows the transcript of the first half of the recording (up to 60 seconds), the first half of its words (up to 100) with times, and the length of the whole recording. Until you unlock it, the result can’t be downloaded or copied. A Pro or Premium pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.
Is my recording uploaded?
No. It is decoded and transcribed in your browser on your device; only the speech model is downloaded, once, from MySmartCoPilot, after you agree. Closing the tab removes the recording and the transcript from the page.
How accurate is it?
For one clear speaker in a widely spoken language most words come out right, with punctuation. Noise, music, accents, crosstalk, names and rare words cause mistakes, more so with the smaller models. Most accurate (Whisper small) is clearly better for difficult recordings and for Indian languages; it is slower.
Which languages does it understand?
The 99 languages Whisper was trained on, including English, Hindi, Bengali, Tamil, Telugu, Marathi, Urdu, Gujarati, Kannada, Malayalam, Punjabi, Spanish, French, German, Portuguese, Arabic, Chinese and Japanese. Moonshine (Fast English) understands English only.
What does the free preview include?
Without a pass the page transcribes the first half of your recording, up to 60 seconds (ending at a pause, so no word is cut), and shows the first half of its words, up to 100, with their times, together with the length of the whole recording. Downloading, copying and making subtitles need the full transcript: a Pro pass, or an unlock of this one result, transcribes the rest on your device from the file still on the page.
Can I correct mistakes in the transcript?
Yes, once you have the full transcript: download it as a Word document, as text or as SRT or WebVTT subtitles, and correct it in the program you use for them.
Does it work offline?
Once a model has been downloaded, the transcription itself needs no connection; the page needs one to open and to download a model the first time.