Speech to text
Dictate with your microphone and watch the text appear as you speak, or transcribe a recording, a voice memo, or a video.
Files are transcribed on your device. Live dictation is handled by your browser, which may use Google’s or Microsoft’s service.
Your browser doesn’t include voice dictation
Firefox and some other browsers don’t yet support speech recognition on web pages. Open this page in Chrome, Edge, or Safari, on a computer or a phone. Meanwhile, you can type or paste text in the box below.
Dictate
Ready to dictate. Select the button and start speaking.
How your voice is processed
This page doesn’t receive or store your voice or your text. Your browser does the recognition: in Chrome and Edge the audio is sent to Google’s or Microsoft’s service to turn it into text, and in Safari Apple processes it.
Live dictation isn’t available in this browser
You can transcribe an audio or video file below. To dictate with the microphone, open this page in Chrome, Edge, or Safari.
Your browser can’t transcribe files
Transcription uses WebAssembly and Web Workers. Open this page in the latest version of Chrome, Edge, Firefox, or Safari.
File
Drop your file hereTap to choose your file
or click to chooseMP3, WAV, M4A, OGG, MP4, WEBM, and more · up to 60 min and 300 MB
Processed on your device: nothing is uploaded
Progress
Transcript
Partial transcript: it stopped before reaching the end of the audio.
Timed segments
Select a time to hear that moment. The .srt and .vtt subtitles use the segments exactly as they came out of the recognizer.
Your audio never leaves your device
Transcription happens in your browser. The only things downloaded, the first time, are the Whisper model from Hugging Face and the engine from jsDelivr; no audio or text is sent anywhere.
How to use it
- Choose “Live dictation” to speak into your microphone, or “Audio file” to transcribe a recording or a video.
- For live dictation, select “Start dictating,” allow the microphone, and speak. Say “comma,” “period,” or “new line” to punctuate.
- For a file, drag it in, choose the quality and the audio language, and select “Transcribe.” The first time, the recognition model is downloaded.
- Edit the text in the box, then copy it or download it as a .txt file. If you transcribed a file, you can also download .srt or .vtt subtitles.
Frequently asked questions
Is my voice or audio sent to a server?
It depends on the mode. For live dictation, your browser does the recognition: in Chrome and Edge the audio goes to Google’s or Microsoft’s service, and in Safari Apple processes it. This page doesn’t receive or store your voice. Files are transcribed on your device with a model that is downloaded once: the audio is never uploaded.
Why doesn’t live dictation work in Firefox?
Firefox doesn’t yet include the speech recognition that web pages use. Open this page in Chrome, Edge, or Safari to dictate. In Firefox you can still use “Audio file”: record your voice with the Voice recorder and transcribe the file.
How do I add periods, commas, and question marks while dictating?
Say them out loud: “period,” “comma,” “colon,” “semicolon,” “question mark,” “exclamation point,” “new line,” “new paragraph,” “open parenthesis,” and “close parenthesis.” The first letter of each sentence is capitalized for you. If you’d rather have the words come out as spoken, turn off “Voice punctuation commands.”
Which files can I transcribe, and how long does it take?
MP3, WAV, M4A, OGG, WEBM, MP4, and other audio or video formats your browser can read, up to 60 minutes and 300 MB. The time depends on your device: the Fast model usually takes less than the length of the audio on a recent computer, and the Accurate model takes longer. While it works you’ll see the progress and the time left, and you can cancel.
Why does it download a model, and how big is it?
To transcribe without uploading your audio, a recognition model (Whisper) has to run in your browser. The first time, about 60 MB are downloaded with the Fast model or about 95 MB with the Accurate one, engine included. After that they stay saved in your browser and aren’t downloaded again. The Accurate model handles accents, proper names, and quiet voices better.
My browser doesn’t ask for microphone permission. What can I do?
If you blocked it before, select the padlock or the settings icon next to the page address, allow the microphone, and try again. On a phone, also check the browser’s permissions in the system settings. If dictation stops by itself after a silence, the tool turns it back on and the text you already dictated isn’t lost.
Two ways to turn speech into text
Live dictation: you speak and the text appears right away. It works best with a nearby microphone, little background noise, and complete sentences. You don’t need to pause between words: recognition understands whole sentences better than single words.
Audio file: you choose a recording, a voice memo, or a video and the tool transcribes it with the Whisper model, without sending it to any server. It’s useful for interviews, classes, meetings, and voice messages. You can play any segment by selecting its time, edit the text, and download .srt or .vtt subtitles. To make a recording first, use the Voice recorder, and to pull the sound out of a video, use Extract audio from video.
Punctuation commands for dictation
- Marks: “period” or “full stop,” “comma,” “colon,” “semicolon,” “ellipsis.”
- Questions and exclamations: “question mark” and “exclamation point.”
- Structure: “new line” and “new paragraph.”
- Parentheses and quotes: “open parenthesis,” “close parenthesis,” “open quote,” “close quote.”
If your text needs the actual word “period” or “comma,” turn off the punctuation commands. They are available for English and Spanish; for other languages the words come out as the browser hears them.
Tips for a clean transcript
- Choose the right language. The dictation language and the audio language of a file are set separately. Pick the variety you speak, such as English (United States) or English (United Kingdom), or let the tool detect the language of a file.
- Check the text afterward. Recognition makes mistakes with names, numbers, and technical terms. Fix them in the box, then clean the result with the Text cleaner or count it with the Word counter.
- Use the Accurate model for hard audio. It takes longer but copes better with accents and quiet voices.
Privacy and credits
This page doesn’t receive your voice or your text. Live dictation uses your browser’s recognition, which usually sends the audio to its maker’s servers. To transcribe files we use OpenAI’s Whisper models (MIT license) with Transformers.js (Apache-2.0) and ONNX Runtime Web (MIT); they are downloaded from Hugging Face and jsDelivr the first time and stay in your browser.
Updated on September 30, 2026