Whisper Web — Free Speech to Text in Your Browser
Whisper Web turns a video or audio file into text without uploading it anywhere. Free Whisper transcription, no sign-up, and subtitles come with it.
- Transcribe a video with Whisper
- Make Whisper SRT subtitles
- MP3 to text
- Whisper without an API key
- Translate speech into English
or drop one here
Choose the language spoken in the recording before you start.
Runs entirely in your browser — your audio is never uploaded. Only the model is downloaded, once.
Whisper is the speech recognition model OpenAI released for anyone to use free of charge. It listens to a recording and writes down what was said, in any of 99 languages, and the Whisper model is good enough that plenty of paid transcription services are quietly built on top of it. Normally, though, getting at Whisper means a developer account, a paid plan or a program to install first.
Whisper Web skips all of that. The Whisper model is loaded into your browser and runs on your own computer, so transcription comes down to picking a file and pressing Transcribe. Your recording never leaves your device — Whisper Web sends nothing to us, nothing to OpenAI and nothing to anyone else. The only thing that travels over the internet is the Whisper model itself, which your browser keeps after the first time, so later runs start straight away.
Whisper Web writes the text on screen as it goes. Copy the Whisper transcript, or download it as a plain text file, as SRT or VTT subtitles for a video editor or a website, or as JSON or a spreadsheet if you want the timings too.
What Whisper Web Does in Your Browser
The real Whisper model, running free
This is OpenAI's own Whisper, not a lighter imitation of it — the same model sitting behind a lot of paid transcription tools. Add an MP4 or an MP3 and Whisper Web puts the text on the page line by line, each with the time it was spoken.
Whisper with no sign-up, no app and no API key
Nothing to create an account for, nothing to install, nothing charged by the minute. Run one file or fifty through Whisper Web: there is no daily limit, no credits to run out of and no watermark on the transcript. Paid services charge because the model runs on their computers — here it runs on yours.
Whisper subtitles, plain text or data
Whisper notes when each line was spoken, so one run gives you five files. SRT for Premiere Pro, DaVinci Resolve, CapCut or YouTube captions; VTT for a video player on a website; TXT to read and search; JSON or CSV if you want the timings in a script or a spreadsheet. Switching between them is instant — nothing is transcribed twice.
Whisper speech to text in 99 languages
Spanish, Mandarin, Arabic, Hindi, Japanese, Ukrainian, Tagalog and plenty more. Tell it which language the recording is in — search the list by name or by code — and it writes that language down. Or switch the task to Translate, and it writes English out of a non-English recording in one go.
Two Whisper models: accuracy, or a smaller download
Turbo is Whisper large-v3 with a trimmed-down decoder, so it transcribes like the largest model without decoding like one — the right choice for lectures, strong accents, people talking over each other and background noise. Base is a fraction of the download and starts sooner, which is enough for a clear voice memo. Each option shows its size, because the model is fetched in full before anything begins.
Whisper Web takes almost any video or audio file
MP4, MOV, WebM and MKV video, or MP3, WAV, M4A, FLAC and OGG audio. Give Whisper Web a video and the sound is taken out for you. The file is read straight off your own disk, so there is no upload bar to sit through — because there is no upload.
Who Uses Whisper Web Transcription
Podcasters writing show notes with Whisper Web
An hour-long episode becomes text you can search, which is where show notes, chapter markers, quote cards and a blog version all start. Running your whole back catalogue through it costs nothing.
Video editors who need Whisper subtitles
Download the SRT that Whisper Web makes and drop it on your timeline. It opens in Premiere Pro, Final Cut, DaVinci Resolve and CapCut like any other subtitle file — no plugin, no waiting on a captioning service, no watermark burned into the picture.
Journalists who need private Whisper transcription
Interview recordings are often the ones you are least free to hand to another company. Whisper Web gives you a first draft to quote and work from, and the audio stays on your own laptop.
Students turning lectures into notes
A recorded lecture or seminar becomes text you can search for the one definition you half remember. Whisper handles technical words and accents far better than a phone's dictation, and there is no student subscription to justify.
Anyone adding captions with Whisper Web
Video should be captioned so people who are deaf or hard of hearing — or just watching with the sound off — can follow it. Make the Whisper subtitle file here and attach it to your video, and an internal training recording never has to be handed to an outside company.
Recordings that cannot be uploaded at all
Medical dictation, legal recordings, HR interviews, unreleased material, anything under an NDA. Browser-based Whisper is often the only route that gets approved, because there is no transfer to approve.
Why Choose Whisper Web
Whisper runs where your recording already is
None of your audio is sent anywhere. The Whisper model comes down to you; your recording does not go up. Close the tab and the transcript is gone, because no copy of it was ever made anywhere else.
Free Whisper, nothing to sign up for
No account, no card, no trial minutes to use up. Whisper Web can be unlimited because the work happens on your computer rather than on ours.
Honest about what Whisper cannot do here
A browser cannot run Whisper's very largest sizes, so a really difficult recording may still come out better from a desktop program or a paid service. It is free, private and instant to start — and for clear speech it is plenty.
How to Transcribe a Video With Whisper Web
1. Add your video or audio file to Whisper Web
Drag in an MP4, MOV, MKV, WebM, MP3, WAV or M4A, or click to browse for one. It is read straight from your computer, so there is no upload to wait for.
2. Choose your Whisper settings
Turbo is selected already, and it is the one to keep if you care about the transcript being right; switch to Base if you would rather not wait for the larger download. Keep Transcribe to write the recording down in its own language, or choose Translate to get English out of a non-English one. Then pick the spoken language — that one is required, and the search box above the list saves scrolling through ninety-nine of them.
3. Press Transcribe and watch the Whisper transcript arrive
The first time, the Whisper model has to download — that is the slow part, and it only happens once. Then the lines appear with their timings as it works, with a progress percentage and a cancel button.
4. Copy the text or download Whisper subtitles
Copy the whole transcript in one click, or pick TXT, SRT, VTT, JSON or CSV and download it. Changing format is instant — nothing has to go through Whisper again.
Whisper Web FAQ
Is Whisper Web really free?
Yes. No account, no card, no credits, no watermark, and no limit on how many files you transcribe or how long they run. It works here because the model runs on your own computer instead of on rented servers — that server time is what every paid Whisper service is charging you for.
Does Whisper Web upload my audio anywhere?
No. Your recording is opened and transcribed inside your browser and is never sent anywhere — not to this site, not to OpenAI, not to anyone. The one thing that does travel over the internet is the Whisper model, which downloads to you and is kept by your browser. After that first download, Whisper Web even works offline.
Do I need an OpenAI API key to use Whisper?
No. Most "Whisper online" pages are really a front door to OpenAI's paid service and need a key first. The model itself is free to download and run, which is what Whisper Web does — no key, no billing account, no per-minute charge.
How accurate is Whisper Web?
At the same size it gives the same result as running Whisper on your desktop, because it is the same model. What changes the accuracy is the size you pick. The two biggest sizes are too large for a browser to handle, so a very noisy or difficult recording may still do better in a desktop program.
Which Whisper model should I choose?
Turbo unless you have a reason not to. It uses the large-v3 encoder, so strong accents, crosstalk and background noise come out far better, and its trimmed decoder keeps it quick. Base is there for when the download matters more than the last few percent of accuracy, or the speech is short and clear. A machine with WebGPU gets through Turbo comfortably; without it the model runs on the CPU, where Base is the more realistic choice for anything long.
Which languages does Whisper handle?
All 99 that Whisper knows, from Spanish, Mandarin and Arabic through to Welsh, Yoruba and Māori. You choose which one before the run starts; there is no automatic detection, because Whisper decides per thirty-second window rather than once per recording, and a long file it finds ambiguous ends up switching language partway through. Translation only goes one way — into English — which is a limit of Whisper itself, not of this page.
Can Whisper Web make SRT or VTT subtitles?
Yes. Every Whisper line is timed, so SRT, VTT, TXT, JSON and CSV are all one click away and switching between them is instant. The SRT opens in Premiere Pro, DaVinci Resolve, CapCut and YouTube's caption uploader like any other subtitle file.
How long a recording can Whisper Web transcribe?
Up to an hour at a time. That is not a rule we invented: a browser tab can only hold so much audio at once, and an hour already fills a good deal of it. For anything longer, cut it into parts with the trim tool on this site and run them through Whisper one after another.
Why is the first Whisper Web run so slow?
The Whisper model has to arrive first — roughly 120MB, 206MB or 586MB depending on the size you picked. Your browser keeps it, so every later run with that size starts immediately. Chrome, Edge and Opera can also put your graphics card to work, which makes the transcribing much faster; Safari and Firefox take longer, but the transcript is identical.
Can I use Whisper transcripts for work or commercially?
Yes. No rights are claimed over your audio or over the text that comes out of it, and nothing is added to either. Whisper is released by OpenAI under an open licence, and Whisper Web is built on the open-source whisper-web project.
Need a link rather than a download? Upload the finished file and get a watch page, a direct file URL and an embed snippet.
Turn a video into a URL