Nari Labs Dia / Dia2
Apache-2.0 dialogue TTS that writes two-speaker conversations with non-verbal cues. Dia2 (1B/2B, Nov 2025) adds streaming; English only, up to 2 minutes.
Overview
Best for: Podcast-style two-speaker dialogue generation.
At a glance
Dia2-2B (also 1B). Two-speaker dialogue, up to 2 minutes of generation.
Script with [S1]/[S2] speaker tags
Waveform (Mimi codec for Dia2)
English
Speaker conditioning/cloning from prompt audio
No numeric claim; Dia2 server described as real streaming.
Wherever you deploy it
Your own deployment; no vendor data processing
GPU
Apache-2.0 (third-party assets such as Kyutai Mimi keep their own licences)
Features
- multi-speaker dialogue
- non-verbal sounds
- streaming server (Dia2)
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- Clone github.com/nari-labs/dia2 and follow the README (GPU).
- Run the Dia2 TTS server for streaming.
Endpoint
Local
Authentication
None
Warnings
2-minute cap
Dia2 generates up to 2 minutes; long content must be segmented.
Identity misuse
Nari Labs explicitly prohibits producing audio resembling real people without permission.
Dialogue-first
Optimised for two-speaker scripts, not single-voice agent replies.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Up to 2 minutes of generation
- English only
Models and products
| Name | Status |
|---|---|
| nari-labs/Dia2-2B, Dia2-1B | Released 2025-11/12 |
| nari-labs/Dia-1.6B-0626 | Older |
Docs and sources
Docs
Sources used
Latency and hardware minimums.