
★ 19k
November 2025
Apache-2.0
Voice & audio
What it does
A 1.6B text-to-speech model that renders an entire scripted conversation at once rather than stitching separate takes, so pacing and emotional flow hold together, and it can produce nonverbals — laughter, a cough, a throat clear. Conditioning on reference audio steers tone and voice, which is also the reason to be careful whose voice you feed it. English only at present, and a successor (Dia2) has since been released. Apache-2.0 for the code.
Summary based on the project's own description. The Japanese page has a fuller beginner guide: setup difficulty, what to prepare, and a step-by-step start.
Screenshots & visuals

Before you start
- Setup difficultyEasy to start
- Command lineRequired
- API keyNot required
- Sends data onlineNo
Judged automatically from the project's own metadata. Always confirm on the official page.
dia FAQ
- Is dia free?
- Yes. It runs on your own machine, so there is no usage-based cost either.
- Do I need to install anything or write code to use dia?
- It needs the command line. If you are not comfortable typing commands, expect to get stuck during setup.
- Does dia work on a phone?
- No. It needs a computer.
- Can I use dia commercially?
- The licence is Apache-2.0, which permits commercial use. Conditions such as keeping the copyright notice still apply, so read the licence text before shipping it inside something.
How it compares
The conditions that decide whether you can actually use it, side by side.
| Tool | Command line | API key | Sends data out | Phone | Commercial use |
|---|---|---|---|---|---|
| dia (this page) | Required | Not needed | No | No | Yes |
| VibeVoice-ComfyUI | Required | Not needed | No | No | Yes |
| Irodori-TTS | Required | Not needed | Yes | No | Yes |
| Chatterbox-TTS-Extended | Required | Not needed | Yes | No | Yes |
"Sends data out" means your input is uploaded to someone's server. Pick "No" when the material cannot leave your organisation.