GA Nari Labs

Nari Labs Dia / Dia2

Apache-2.0 dialogue TTS that writes two-speaker conversations with non-verbal cues. Dia2 (1B/2B, Nov 2025) adds streaming; English only, up to 2 minutes.

Est. per minuten/a

Overview

Best for: Podcast-style two-speaker dialogue generation.

At a glance

TypeText-to-speech
Params B2
LicenceApache-2.0
CommercialYes
StreamingYes
Languages1

Dia2-2B (also 1B). Two-speaker dialogue, up to 2 minutes of generation.

Audio in

Script with [S1]/[S2] speaker tags

Audio out

Waveform (Mimi codec for Dia2)

Languages

English

Voices

Speaker conditioning/cloning from prompt audio

Latency

No numeric claim; Dia2 server described as real streaming.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

GPU

Licence

Apache-2.0 (third-party assets such as Kyutai Mimi keep their own licences)

Features

  • multi-speaker dialogue
  • non-verbal sounds
  • streaming server (Dia2)

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. Clone github.com/nari-labs/dia2 and follow the README (GPU).
  2. Run the Dia2 TTS server for streaming.

Endpoint

Local

Authentication

None

Warnings

2-minute cap

Dia2 generates up to 2 minutes; long content must be segmented.

Identity misuse

Nari Labs explicitly prohibits producing audio resembling real people without permission.

Dialogue-first

Optimised for two-speaker scripts, not single-voice agent replies.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Up to 2 minutes of generation
  • English only

Models and products

NameStatusNotes
nari-labs/Dia2-2B, Dia2-1BReleased 2025-11/12Streaming dialogue TTS.
nari-labs/Dia-1.6B-0626Older

Docs and sources

Docs

Sources used

Not fully verified

Latency and hardware minimums.

Similar open models APIs

Spotted a wrong price or a dead link?