Deepgram Aura-2
Enterprise-focused TTS tuned for clear, business-style voice agents. WebSocket input streaming with Speak/Flush/Clear/Close messages; generous concurrency on pay-as-you-go.
Overview
Best for: Enterprise voice agents already on Deepgram STT; high concurrency on pay-as-you-go; self-hosting needs.
At a glance
Sub-200 ms vendor claim. $30/1M PAYG, $27/1M Growth; Aura-1 is half price. Concurrency 45 on PAYG (REST and WSS combined). Socket max 60 minutes, 2,400 chars per minute per socket, 20 flushes per minute. HIPAA and SOC 2 are Deepgram vendor claims, not re-verified.
Plain text
WebSocket: linear16, mulaw, alaw (container none); linear16 8/16/24/32/48 kHz, mulaw/alaw 8 or 16 kHz. REST adds mp3 (22.05 kHz), opus (48 kHz ogg), flac, aac.
7 for Aura-2: en, es, nl, fr, de, it, ja (changelog Dec 2025 / Jan 2026)
40+ English voices per Deepgram marketing; ~90 total across languages per a third-party catalog. Voice cloning: not offered on the pages checked.
Vendor claim: sub-200 ms (Aura-2 product page). No figure on the streaming docs page.
Hosted API plus self-hosted/on-prem option
Self-hosted deployment available (Deepgram docs). Certifications not re-verified here.
Features
- input streaming (Speak + Flush)
- Clear message for barge-in
- self-hosted deployment option
- telephony encodings
Pricing
| What | Price | Unit |
|---|---|---|
| Aura-2 | $0.030 | per 1K characters |
| Aura-2 | $0.027 | per 1K characters |
| Aura-1 | $0.0150 / $0.0135 | per 1K characters |
900 chars/min; Aura-2 Growth vs PAYG
Free tier: $200 credit for new accounts
Source: deepgram.com
Setup
- Create an API key in the Deepgram console.
- Connect to wss://api.deepgram.com/v1/speak with model, encoding and sample_rate query params.
- Send {type:'Speak', text} messages as tokens arrive, then {type:'Flush'} at end of turn.
- Write binary frames as audio; JSON frames are control/metadata (Flushed, Warning).
Endpoint
wss://api.deepgram.com/v1/speak
Authentication
Authorization: Token <API_KEY> header
Quick start javascript
// npm i ws
import WebSocket from "ws";
import fs from "fs";
const url = "wss://api.deepgram.com/v1/speak?model=aura-2-thalia-en&encoding=linear16&sample_rate=24000";
const ws = new WebSocket(url, { headers: { Authorization: `Token ${process.env.DEEPGRAM_API_KEY}` } });
const out = fs.createWriteStream("out_24k_s16le.pcm");
ws.on("open", () => {
for (const t of ["Hello there. ", "This text is streamed ", "sentence by sentence."]) {
ws.send(JSON.stringify({ type: "Speak", text: t }));
}
ws.send(JSON.stringify({ type: "Flush" }));
});
ws.on("message", (data, isBinary) => {
if (isBinary) return out.write(data);
const msg = JSON.parse(data.toString());
if (msg.type === "Flushed") ws.send(JSON.stringify({ type: "Close" }));
});
ws.on("close", () => out.end());
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Flush rate limit
Only 20 Flush messages per 60 s per socket. Flushing after every LLM token or clause will trigger warnings; flush once per turn and let sentence punctuation drive synthesis.
Streaming formats are limited
The WebSocket only outputs linear16, mulaw and alaw. MP3/Opus are REST-only, so browsers need a PCM player or you transcode.
One voice per connection
Model/voice and encoding are fixed at connect time. Switching voice mid-call means a new socket; Deepgram recommends one socket per conversation.
Throughput cap per socket
2,400 characters per minute per socket is fine for one live speaker but too slow for bulk narration; use REST for long-form.
WAV headers cause clicks in telephony
For REST telephony output set container=none; WAV headers mid-stream produce audible clicks.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- 2,000 characters per Speak payload (413 above that)
- 2,400 characters per minute throughput per socket
- 20 Flush messages per 60 s
- 60-minute max socket lifetime
- Voice and output settings fixed per connection
- TTS concurrency: PAYG 45, Growth 60 (REST + WSS combined)
Models and products
| Name | Status |
|---|---|
| aura-2-<voice>-<lang> (e.g. aura-2-thalia-en) | GA |
| aura-<voice>-en (Aura-1) | GA (older) |
Docs and sources
Docs
Sources used
- deepgram.com/pricing
- developers.deepgram.com/docs/streaming-text-to-speech
- developers.deepgram.com/docs/tts-media-output-settings
- deepgram.com/learn/aura-2-now-speaks-dutch-french-german-italian-japanese
Exact voice count; latency is vendor marketing only.