Deepgram Voice Agent API
One WebSocket that runs Deepgram speech recognition, a managed or bring-your-own LLM, and Deepgram Aura or third-party TTS, billed per connected minute. Good for developers who want a simple flat rate with the LLM included on Standard models.
Overview
Best for: Developers who want one WebSocket, predictable per-minute cost and strong telephony-grade STT, with an option to swap in their own LLM or TTS.
At a glance
Price is Standard tier PAYG (BYO LLM+TTS $0.050, Advanced $0.163). Concurrency 45 on PAYG, 60 on Growth. English by default, multilingual via multi models. Deepgram markets SOC 2 and HIPAA readiness, not verified.
linear16 default at 16 kHz (other encodings/rates configurable)
linear16 (e.g. 24 kHz), optional container; other encodings configurable
English by default; multilingual via Nova 'multi' or flux-general-multi with language hints, and multilingual TTS providers.
Deepgram Aura-2 voices (e.g. aura-2-thalia-en) or third-party voices.
Not stated on the pages read; the API reports per-turn latency metrics (time to first token, TTS time to first byte).
Hosted API (agent.deepgram.com); Deepgram also sells self-hosted deployments (not verified for the agent API).
Not re-verified this session (Deepgram markets SOC 2 and HIPAA readiness; check the trust page).
Features
- function calling (client or server side)
- barge-in (UserStartedSpeaking)
- greeting
- LLM fallback
- reusable stored agent configurations
- browser agent SDK and widget
- mid-session updates of prompt/think/speak
Pricing
| What | Price | Unit |
|---|---|---|
| Standard | $0.075 PAYG / $0.068 Growth | per minute |
| Standard - BYO TTS | $0.065 / $0.051 | per minute |
| Custom - BYO LLM | $0.065 / $0.059 | per minute |
| Custom - BYO LLM + TTS | $0.050 / $0.041 | per minute |
| Advanced | $0.163 / $0.146 | per minute |
| Advanced - BYO TTS | $0.122 / $0.110 | per minute |
Official tier rates; BYO tiers add your own LLM/TTS bills on top.
Not token-billed; billed on WebSocket connection time.
Free tier: $200 one-time credit for new accounts.
Source: deepgram.com
Setup
- Create a Deepgram account (includes $200 credit) and an API key.
- Open wss://agent.deepgram.com/v1/agent/converse with Authorization: Token <key>.
- Send a Settings message (audio formats, listen/think/speak providers, prompt, greeting).
- Stream raw audio as binary frames, play binary audio back, send KeepAlive when idle, stop playback on UserStartedSpeaking.
Endpoint
wss://agent.deepgram.com/v1/agent/converse
Authentication
Authorization header (Token <API key>, or a short-lived token for browsers)
Quick start javascript
// npm i ws
import WebSocket from "ws";
const ws = new WebSocket("wss://agent.deepgram.com/v1/agent/converse", {
headers: { Authorization: `Token ${process.env.DEEPGRAM_API_KEY}` },
});
ws.on("open", () => {
ws.send(JSON.stringify({
type: "Settings",
audio: {
input: { encoding: "linear16", sample_rate: 16000 },
output: { encoding: "linear16", sample_rate: 24000, container: "none" },
},
agent: {
listen: { provider: { type: "deepgram", model: "flux-general-en", version: "v2" } },
think: { provider: { type: "open_ai", model: "gpt-4o-mini" }, // a "Standard" tier model
prompt: "You are a helpful assistant. Keep answers short." },
speak: { provider: { type: "deepgram", model: "aura-2-thalia-en" } },
greeting: "Hi! How can I help?",
},
}));
setInterval(() => ws.send(JSON.stringify({ type: "KeepAlive" })), 5000);
});
// Send raw linear16 mic audio as binary frames
export function sendPcm(buf) { ws.send(buf); }
ws.on("message", (data, isBinary) => {
if (isBinary) return play(data); // agent speech (raw PCM)
const ev = JSON.parse(data.toString());
if (ev.type === "UserStartedSpeaking") stopPlayback(); // barge-in
if (ev.type === "Error") console.error(ev);
});
function play(pcm) {} function stopPlayback() {}
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Model choice changes the tier
Picking an Advanced-tier LLM (e.g. GPT-5, Claude Sonnet) moves the session to $0.163/min, more than double Standard. The tier follows the model, so a config change can double cost.
Connection time is billed
The meter runs while the WebSocket is open, including silence and hold. Close sockets promptly and avoid idle KeepAlive loops.
BYO tiers are not all-in
BYO LLM/TTS rates exclude what you pay OpenAI, Bedrock, Groq, ElevenLabs etc. Bedrock and Groq always require your own endpoint and credentials.
Promotional rates expired
Third-party pages still quote promo or older rates (e.g. $0.08 Standard, $0.056 BYO LLM); some promos ended 2026-09-12. Use deepgram.com/pricing.
Concurrency ceiling
45 concurrent sessions on PAYG and 60 on Growth; larger contact centres need Enterprise.
Deprecated fields
agent.language is deprecated; set language on listen.provider and speak.provider. Some listed models are already marked deprecated.
Cascaded, not native speech-to-speech
STT, LLM and TTS are separate stages, so paralinguistic cues (tone, laughter) are not passed to the LLM.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Concurrency: up to 45 WebSocket agent sessions on Pay As You Go, up to 60 on Growth
- Billing runs for the whole WebSocket connection
Models and products
| Name | Status |
|---|---|
| STT: Deepgram Flux / Nova (e.g. flux-general-en, flux-general-multi) | GA |
| LLM Standard tier | GA |
| LLM Advanced tier | GA |
| TTS: Deepgram Aura-2 / Flux TTS, or eleven_labs, cartesia, open_ai, aws_polly | GA |
Docs and sources
Docs
Sources used
- deepgram.com/pricing
- developers.deepgram.com/docs/voice-agent-llm-models.md
- developers.deepgram.com/docs/configure-voice-agent.md
- developers.deepgram.com/reference/voice-agent/voice-agent.md
Latency figures; compliance certifications; Growth plan entry requirements; whether 'Token' vs 'Bearer' is required for every key type.