Text to SpeechBeta

Type something and hear it spoken by a natural-sounding AI voice running on your own device — then save it as a WAV file.

Type something and hear it spoken by a natural-sounding voice model running on your device — nine English speakers, plus four other languages. Save the result as a WAV file. The text is never sent anywhere.

You need a recent browser, some free disk space, and patience on the first download. A GPU (WebGPU) makes everything many times faster and is used automatically when this device passes a real GPU test; without one, everything runs on the CPU.

Choose a model and press Download & load to begin.

The first load downloads the weights into this browser's storage. Later visits reuse them and start in seconds.

Resource usage no model loaded
CPU load 0% busy
CPU threads
Memory
GPU
Model storage
Output speed

Browsers do not report CPU or GPU load to a web page, so nothing here is invented. CPU load is measured two ways and shows the higher: how much of each second the engine spends blocking its own worker, and how much slower a fixed piece of arithmetic runs right now than on an idle machine — which is what a multi-threaded engine looks like from here. Memory is the JavaScript and WebAssembly heap, which only Chromium-based browsers expose. GPU utilisation and video memory are not available in any browser.

Generation settings

Create an image

No model loaded yet.

Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.

Transcribe audio

No model loaded yet.
No audio chosen yet.

Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.

Read text aloud

No model loaded yet.

Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.

Make music

No model loaded yet.

MusicGen is slow without a GPU — measured on a plain CPU, about 50 seconds of computing for 2 seconds of music. Start with a short clip.

Speech to text, text to speech, music and images share one model slot, so loading one replaces the previous one — your chat model stays where it is.

Load any model from Hugging Face

Search the Hub, or paste a repository id. Each weight variant is listed with its real download size and whether it is likely to fit on this device.

Downloaded on this device

Models are stored by your browser and count towards its storage quota. Delete any of them here at any time.

Saved conversations

Your chats are kept in this browser only, so you can pick one up on your next visit. They are never uploaded, and deleting models above does not touch them.

This device

Graphics drivers and browsers get updated. If GPU inference was switched off after a failure, you can ask for it to be tested again.

Log


  

🔒 Runs entirely in your browser — nothing you type is uploaded or stored on a server.

Text to Speech reads your text aloud with a voice model that runs inside your browser. The English voices come from Kokoro, an open 82-million-parameter model that sounds far closer to a person than the robotic read-aloud built into most computers. You get its nine best-rated speakers, American and British, male and female, and a speed control. It is a 92 MB download, once; after that it works offline.

Also on Txtset: write the script with a local AI model first · transcribe a recording back into text · add a music bed under the voice · make a picture to go with it.

The result is a real audio file, not just playback: press Download WAV and you have a clip you can drop into a video editor, a slide deck or a lesson. Because nothing is sent to a server, it is safe for scripts, drafts and anything else you would rather not paste into a cloud service.

Hearing your own writing read back is also the quickest way to proofread it — a missing word or a clumsy sentence is obvious to the ear in a way it never is on screen. For Spanish, French, German and Russian there are also Meta's MMS voices, one small 38 MB model per language — clear but flatter than Kokoro, and licensed for non-commercial use only.

Be realistic about speed on an ordinary computer: measured on a desktop processor with no graphics card, Kokoro speaks about 7 seconds of audio in 20 seconds of work, sentence by sentence, with a progress bar. A paragraph takes moments; a whole chapter is better done in sections.

How to use

  1. Pick a model and press Download & load — Kokoro for English (92 MB), or an MMS voice for another language (38 MB).
  2. Choose a speaker and a speed, then type or paste your text into the box.
  3. Press Speak to generate and play it.
  4. Press Download WAV to save the audio.

Examples

Proofread by ear
Paste a draft and listen — you will catch mistakes your eyes skip over.
Quick voice-over
Generate narration for a short explainer video or a slide.
Language practice
Hear how a Spanish, French or German sentence sounds.
Accessibility
Have a long passage read aloud when reading on screen is tiring.

Frequently asked questions

Is it really free?
Yes. There is no account, no character limit and no credit system, because the voice runs on your own device — there is no server cost to meter.
Is my text uploaded?
No. The text goes to the voice model inside this tab and nowhere else. The only download is the model itself.
Can I download the audio?
Yes — every clip can be saved as a WAV file, which any audio or video editor opens.
Can I use the audio commercially?
With the Kokoro English voices, yes: the model is published under the Apache 2.0 licence, which permits commercial use. The MMS voices for other languages are different — Meta publishes them under CC-BY-NC 4.0, which rules out commercial use, so keep those to proofreading, study and personal projects.
Can I choose a male or female voice?
Yes, in English. Kokoro offers American and British speakers, male and female — the list shows the best-rated ones — plus a speed control. The MMS voices for other languages have one speaker each.
How long a text can it read?
Paragraphs are fine. Very long texts take longer to synthesise and make a large file, so split a long document into sections.
Does it work offline?
Once the voice has been downloaded, yes — it is stored in your browser and loads without a connection.