Speech to TextBeta
Turn a recording into text with Whisper, running inside this tab — the audio never leaves your device.
Turn a recording into text without uploading it. Choose a small speech model, drop in an audio or video file, and Whisper transcribes it inside this tab — in 99 languages, with timestamps if you want them.
You need a recent browser, some free disk space, and patience on the first download. A GPU (WebGPU) makes everything many times faster and is used automatically when this device passes a real GPU test; without one, everything runs on the CPU.
Choose a model and press Download & load to begin.
The first load downloads the weights into this browser's storage. Later visits reuse them and start in seconds.
Resource usage no model loaded
Browsers do not report CPU or GPU load to a web page, so nothing here is invented. CPU load is measured two ways and shows the higher: how much of each second the engine spends blocking its own worker, and how much slower a fixed piece of arithmetic runs right now than on an idle machine — which is what a multi-threaded engine looks like from here. Memory is the JavaScript and WebAssembly heap, which only Chromium-based browsers expose. GPU utilisation and video memory are not available in any browser.
Generation settings
Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.
Transcribe audio
Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.
Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.
Make music
MusicGen is slow without a GPU — measured on a plain CPU, about 50 seconds of computing for 2 seconds of music. Start with a short clip.
Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.
Load any model from Hugging Face
Search the Hub, or paste a repository id. Each weight variant is listed with its real download size and whether it is likely to fit on this device.
Downloaded on this device
Models are stored by your browser and count towards its storage quota. Delete any of them here at any time.
Saved conversations
Your chats are kept in this browser only, so you can pick one up on your next visit. They are never uploaded, and deleting models above does not touch them.
This device
Graphics drivers and browsers get updated. If GPU inference was switched off after a failure, you can ask for it to be tested again.
Log
🔒 Runs entirely in your browser — nothing you type is uploaded or stored on a server.
Speech to Text transcribes an audio or video file inside your browser. The recognition model — OpenAI's Whisper, or Useful Sensors' Moonshine — is downloaded once and then runs on your own device, so an interview, a voice memo or a lecture recording is never uploaded to a transcription service.
Also on Txtset: summarise the transcript with a local AI model · go the other way and read text aloud · generate music on the same device · cut the clip down before you transcribe it.
Choose the model that fits the job. Whisper Tiny (English) is the quickest and a 96 MB download; Whisper Base is 142 MB, understands 99 languages and can translate what it hears straight into English; Moonshine Base is a newer English model that is quick on short clips. Long recordings are cut into overlapping 30-second windows, the way Whisper expects, so nothing at the end is silently dropped.
Tick timestamps to get a line per phrase with the time it starts, which is what you want for subtitles or for finding a quote in a long recording. The transcript lands in a box you can copy, and from there the other tools on this site take over — count it, clean it up, or ask the local chat model to summarise it.
How to use
- Pick a model and press Download & load (96–142 MB, once).
- Choose an audio or video file — WAV, MP3, M4A, OGG and WebM all work; the browser decodes it.
- Leave the language on Detect or pick it, tick timestamps or translate to English if you want them.
- Press Transcribe and copy the text when it appears.
Examples
Turn a recorded interview into text you can quote from, without sending it to a third party.
Get a searchable transcript of a class or a talk, with timestamps to jump back to.
Convert a rambling voice note into text you can edit into shape.
Let Whisper Base translate a clip in another language straight into English text.