GA Deepgram

Deepgram Aura-2

Enterprise-focused TTS tuned for clear, business-style voice agents. WebSocket input streaming with Speak/Flush/Clear/Close messages; generous concurrency on pay-as-you-go.

Est. per minute$0.024 - 0.027
1 high-severity warning

Overview

Best for: Enterprise voice agents already on Deepgram STT; high concurrency on pay-as-you-go; self-hosting needs.

At a glance

$/1M chars$30
Free tierYes
Free credit $$200
Text stream inYes
8 kHz phoneYes
Latency ms200
Languages7
Max session min60
Concurrency45
WebRTCNo
WebSocketYes
gRPCNo
HIPAAYes
SOC 2Yes
Self-hostYes
Open weightsNo

Sub-200 ms vendor claim. $30/1M PAYG, $27/1M Growth; Aura-1 is half price. Concurrency 45 on PAYG (REST and WSS combined). Socket max 60 minutes, 2,400 chars per minute per socket, 20 flushes per minute. HIPAA and SOC 2 are Deepgram vendor claims, not re-verified.

Audio in

Plain text

Audio out

WebSocket: linear16, mulaw, alaw (container none); linear16 8/16/24/32/48 kHz, mulaw/alaw 8 or 16 kHz. REST adds mp3 (22.05 kHz), opus (48 kHz ogg), flac, aac.

Languages

7 for Aura-2: en, es, nl, fr, de, it, ja (changelog Dec 2025 / Jan 2026)

Voices

40+ English voices per Deepgram marketing; ~90 total across languages per a third-party catalog. Voice cloning: not offered on the pages checked.

Latency

Vendor claim: sub-200 ms (Aura-2 product page). No figure on the streaming docs page.

Regions

Hosted API plus self-hosted/on-prem option

Compliance

Self-hosted deployment available (Deepgram docs). Certifications not re-verified here.

Features

  • input streaming (Speak + Flush)
  • Clear message for barge-in
  • self-hosted deployment option
  • telephony encodings

Pricing

WhatPriceUnitNotes
Aura-2$0.030per 1K charactersPay As You Go
Aura-2$0.027per 1K charactersGrowth plan
Aura-1$0.0150 / $0.0135per 1K charactersPAYG / Growth
How the per-minute estimate was worked out

900 chars/min; Aura-2 Growth vs PAYG

Free tier: $200 credit for new accounts

Source: deepgram.com

Setup

  1. Create an API key in the Deepgram console.
  2. Connect to wss://api.deepgram.com/v1/speak with model, encoding and sample_rate query params.
  3. Send {type:'Speak', text} messages as tokens arrive, then {type:'Flush'} at end of turn.
  4. Write binary frames as audio; JSON frames are control/metadata (Flushed, Warning).

Endpoint

wss://api.deepgram.com/v1/speak

Authentication

Authorization: Token <API_KEY> header

Quick start javascript

// npm i ws
import WebSocket from "ws";
import fs from "fs";

const url = "wss://api.deepgram.com/v1/speak?model=aura-2-thalia-en&encoding=linear16&sample_rate=24000";
const ws = new WebSocket(url, { headers: { Authorization: `Token ${process.env.DEEPGRAM_API_KEY}` } });
const out = fs.createWriteStream("out_24k_s16le.pcm");

ws.on("open", () => {
  for (const t of ["Hello there. ", "This text is streamed ", "sentence by sentence."]) {
    ws.send(JSON.stringify({ type: "Speak", text: t }));
  }
  ws.send(JSON.stringify({ type: "Flush" }));
});
ws.on("message", (data, isBinary) => {
  if (isBinary) return out.write(data);
  const msg = JSON.parse(data.toString());
  if (msg.type === "Flushed") ws.send(JSON.stringify({ type: "Close" }));
});
ws.on("close", () => out.end());

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Flush rate limit

Only 20 Flush messages per 60 s per socket. Flushing after every LLM token or clause will trigger warnings; flush once per turn and let sentence punctuation drive synthesis.

Streaming formats are limited

The WebSocket only outputs linear16, mulaw and alaw. MP3/Opus are REST-only, so browsers need a PCM player or you transcode.

One voice per connection

Model/voice and encoding are fixed at connect time. Switching voice mid-call means a new socket; Deepgram recommends one socket per conversation.

Throughput cap per socket

2,400 characters per minute per socket is fine for one live speaker but too slow for bulk narration; use REST for long-form.

WAV headers cause clicks in telephony

For REST telephony output set container=none; WAV headers mid-stream produce audible clicks.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • 2,000 characters per Speak payload (413 above that)
  • 2,400 characters per minute throughput per socket
  • 20 Flush messages per 60 s
  • 60-minute max socket lifetime
  • Voice and output settings fixed per connection
  • TTS concurrency: PAYG 45, Growth 60 (REST + WSS combined)

Models and products

NameStatusNotes
aura-2-<voice>-<lang> (e.g. aura-2-thalia-en)GAModel and voice are one ID. English, Spanish, plus Dutch, French, German, Italian, Japanese added Dec 2025.
aura-<voice>-en (Aura-1)GA (older)Half the price of Aura-2.

Docs and sources

Docs

Sources used

Not fully verified

Exact voice count; latency is vendor marketing only.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?