The Best Open Source Text to Speech Models

Piper is the best choice for fast, lightweight offline speech; Kokoro gives natural quality from a small, permissively licensed model; XTTS, F5-TTS and Chatterbox add voice cloning. Check weight licences carefully, since several cloning models restrict commercial use.
What should you look for in an open source TTS model?
Text to speech models differ on five axes: naturalness, speed, languages, voice cloning and licence. No single project wins all five, so the right pick depends on whether you are building a voice agent, an audiobook pipeline, an accessibility feature or an embedded device.
The licence trap is specific to TTS. Code is often MIT or Apache 2.0, while the trained weights carry a separate, sometimes non-commercial licence. Read the model card of the checkpoint, not only the repository licence.
Voice cloning also raises consent issues. Only clone voices you have explicit permission to use, and label synthetic audio where your users could mistake it for a real person.
Which open source TTS models are worth considering?
| Model | Strength | Licence notes | Best for | Trade-off |
|---|---|---|---|---|
| Piper | Very fast, runs on CPU and Raspberry Pi | Check the licence file and per-voice terms | Home assistants, embedded, offline apps | Less expressive than larger models |
| Kokoro | Natural voices from a small model | Weights released under Apache 2.0 | Voice agents and apps needing quality on modest hardware | Limited set of preset voices, no cloning |
| XTTS (Coqui) | Multilingual voice cloning from short samples | Coqui Public Model License, non-commercial | Research and personal projects | Commercial use not permitted by the model licence |
| F5-TTS | High-quality zero-shot cloning | Code MIT; check weight licence (often non-commercial) | Experiments with expressive cloning | Heavier inference, licence limits |
| Chatterbox | Cloning with emotion control | Check the licence file | Products needing expressive cloned voices | Newer project, verify stability |
| StyleTTS 2 | Natural prosody, trainable | Code MIT; check pretrained model terms | Custom single-speaker voices | Training and setup take effort |
| Bark | Generates speech plus laughs and sounds | MIT | Creative audio and prototypes | Slow and less controllable |
When are Piper or Kokoro the right choice?
Piper uses compact VITS-style voices exported to ONNX and runs faster than real time on ordinary CPUs, including single-board computers. That is why the Home Assistant ecosystem adopted it for local voice control.
It supports many languages through community-trained voices, each with its own quality and terms. The sound is clear and pleasant but less expressive than large modern models, which matters for audiobooks more than for assistants.
Because each voice is a small file, you can ship several voices inside an app without large downloads, which suits offline mobile and kiosk products.
Kokoro delivers speech that sounds close to much larger models while staying small enough for a CPU or a modest GPU. Its weights are published under Apache 2.0, which makes it unusually straightforward for commercial products.
It ships with preset voices rather than cloning. For many products, such as reading articles aloud or giving an agent a consistent voice, a good preset is exactly what you need and avoids consent questions entirely.
Which models support voice cloning?
XTTS, F5-TTS, Chatterbox and several newer projects can imitate a voice from a short reference clip. Quality varies with the sample: clean, noise-free audio of a single speaker gives the best results.
Coqui, the company behind XTTS, has shut down, and its model licence does not allow commercial use, although the community still maintains forks of the code. F5-TTS weights have also been published under non-commercial terms. If cloning is central to a paid product, confirm the weight licence or train your own voice on data you own.
Cloning quality also depends on the reference clip matching the target style. A calm narration sample produces calm output, so record references in the tone you want the final audio to have.
How much hardware does open source TTS need?
- Piper: any modern CPU, including a Raspberry Pi 4 or 5, with real-time output.
- Kokoro: runs on CPU; a small GPU makes it comfortably faster than real time for interactive use.
- Cloning models such as XTTS, F5-TTS and Chatterbox: a GPU with several gigabytes of memory for acceptable speed.
- Bark: GPU strongly recommended; generation is slow compared with dedicated TTS models.
- Batch audiobook generation: any of the above, since latency matters less than total throughput.
How to choose a TTS model step by step
- Decide whether you need preset voices or cloning; presets are simpler legally and technically.
- List target languages and check that each candidate has good voices for them, not just nominal support.
- Set a latency goal: voice agents need the first audio chunk quickly, which favors streaming-capable small models.
- Read the weight licence and any per-voice terms before prototyping, not after.
- Generate the same 10 scripts with every candidate, including numbers, names and abbreviations, and listen blind.
- Plan text normalization, because most failures come from dates, currencies and acronyms rather than the model.
Can you train a custom voice?
Yes, with the right model. Piper voices can be trained or fine-tuned from recordings of a single speaker, and StyleTTS 2 is popular for custom single-voice training. Expect to need clean studio-quality recordings and a GPU for training.
A custom trained voice avoids both licensing doubts about pretrained weights and consent problems, provided the speaker signs off. For brands that want a distinctive voice they own, this is often the cleanest path.
Where open source TTS breaks
- Numbers, units and abbreviations read literally; add a normalization step before synthesis.
- Long paragraphs producing drift in tone; split text into sentences and synthesize them in sequence.
- Mispronounced brand names and foreign words; use phoneme input or a custom lexicon where supported.
- Shipping a non-commercial model in a paid product because the code repository said MIT.
- Cloning voices without documented consent from the speaker.
- Ignoring streaming: waiting for full audio before playback makes voice agents feel slow.
How do you serve TTS in production and inside a voice agent?
Wrap the model in a small HTTP or WebSocket service that accepts text and streams audio back in chunks. Several community projects already expose OpenAI-compatible speech endpoints for models like Kokoro and Piper, which lets existing clients switch with a base URL change.
Cache aggressively. Many products repeat the same phrases, such as greetings, menu options and confirmations, and pre-generated audio costs nothing to serve. Only dynamic text needs live synthesis.
Choose output formats deliberately. Telephony often needs 8 kHz or 16 kHz audio in specific codecs, while web playback prefers compressed formats such as Opus or MP3. Converting at the edge of your service keeps the model code simple.
A typical voice agent chains speech to text, a language model and TTS. The TTS stage decides how responsive the agent feels, so stream the model’s text sentence by sentence into TTS and start playback as soon as the first chunk is ready.
Small models such as Kokoro and Piper suit this loop because they keep time to first audio low on modest hardware. RepoLoot’s catalog tags audio projects by difficulty, which helps you estimate how much integration work a full local voice stack needs.
- Split input into sentences and synthesize them in parallel where order can be preserved.
- Keep the voice, model version and settings in configuration so every environment sounds the same.
- Monitor time to first audio, not only total generation time.
- Log text inputs that produce bad audio and add them to your test scripts.
Frequently asked questions
- What is the best free text to speech model for commercial use?
- Kokoro is a strong choice because its weights are published under Apache 2.0 and quality is high for its size. Piper is another practical option for fast offline speech, but check each voice’s terms. Many cloning models, including XTTS, use non-commercial licences.
- Which open source TTS is closest to ElevenLabs quality?
- Recent models such as Kokoro, F5-TTS and Chatterbox come closest in naturalness, and cloning-capable ones approach the voice-matching features of commercial services. Results vary with language and voice, and licences differ, so compare samples of your own scripts rather than relying on demos.
- Can open source TTS run offline on a Raspberry Pi?
- Yes, Piper was designed for this and runs faster than real time on a Raspberry Pi 4 or 5. It is widely used for local voice assistants. Larger expressive or cloning models generally need a desktop GPU to reach usable speed.
- Is it legal to clone a voice with open source TTS?
- Cloning your own voice or a voice you have explicit, documented permission to use is generally fine, subject to the model’s licence. Cloning someone without consent can violate personality, privacy or fraud laws depending on the country. Get written consent and disclose synthetic audio.