GA Deepgram

Deepgram Voice Agent API

One WebSocket that runs Deepgram speech recognition, a managed or bring-your-own LLM, and Deepgram Aura or third-party TTS, billed per connected minute. Good for developers who want a simple flat rate with the LLM included on Standard models.

Est. per minute$0.041 - 0.163
2 high-severity warnings

Overview

Best for: Developers who want one WebSocket, predictable per-minute cost and strong telephony-grade STT, with an option to swap in their own LLM or TTS.

At a glance

Flat $/min$0.075
Free tierYes
Free credit $$200
Native S2SNo
ToolsYes
Own LLMYes
Concurrency45
WebRTCNo
WebSocketYes
Phone / SIPNo
Open weightsNo

Price is Standard tier PAYG (BYO LLM+TTS $0.050, Advanced $0.163). Concurrency 45 on PAYG, 60 on Growth. English by default, multilingual via multi models. Deepgram markets SOC 2 and HIPAA readiness, not verified.

Audio in

linear16 default at 16 kHz (other encodings/rates configurable)

Audio out

linear16 (e.g. 24 kHz), optional container; other encodings configurable

Languages

English by default; multilingual via Nova 'multi' or flux-general-multi with language hints, and multilingual TTS providers.

Voices

Deepgram Aura-2 voices (e.g. aura-2-thalia-en) or third-party voices.

Latency

Not stated on the pages read; the API reports per-turn latency metrics (time to first token, TTS time to first byte).

Regions

Hosted API (agent.deepgram.com); Deepgram also sells self-hosted deployments (not verified for the agent API).

Compliance

Not re-verified this session (Deepgram markets SOC 2 and HIPAA readiness; check the trust page).

Features

  • function calling (client or server side)
  • barge-in (UserStartedSpeaking)
  • greeting
  • LLM fallback
  • reusable stored agent configurations
  • browser agent SDK and widget
  • mid-session updates of prompt/think/speak

Pricing

WhatPriceUnitNotes
Standard$0.075 PAYG / $0.068 Growthper minuteIncludes Deepgram STT, a Standard-tier managed LLM and Deepgram TTS.
Standard - BYO TTS$0.065 / $0.051per minute
Custom - BYO LLM$0.065 / $0.059per minuteYou also pay your LLM provider.
Custom - BYO LLM + TTS$0.050 / $0.041per minuteYou also pay LLM and TTS providers.
Advanced$0.163 / $0.146per minuteAdvanced-tier LLMs.
Advanced - BYO TTS$0.122 / $0.110per minute
How the per-minute estimate was worked out

Official tier rates; BYO tiers add your own LLM/TTS bills on top.

Audio token rate

Not token-billed; billed on WebSocket connection time.

Free tier: $200 one-time credit for new accounts.

Source: deepgram.com

Setup

  1. Create a Deepgram account (includes $200 credit) and an API key.
  2. Open wss://agent.deepgram.com/v1/agent/converse with Authorization: Token <key>.
  3. Send a Settings message (audio formats, listen/think/speak providers, prompt, greeting).
  4. Stream raw audio as binary frames, play binary audio back, send KeepAlive when idle, stop playback on UserStartedSpeaking.

Endpoint

wss://agent.deepgram.com/v1/agent/converse

Authentication

Authorization header (Token <API key>, or a short-lived token for browsers)

Quick start javascript

// npm i ws
import WebSocket from "ws";

const ws = new WebSocket("wss://agent.deepgram.com/v1/agent/converse", {
  headers: { Authorization: `Token ${process.env.DEEPGRAM_API_KEY}` },
});

ws.on("open", () => {
  ws.send(JSON.stringify({
    type: "Settings",
    audio: {
      input: { encoding: "linear16", sample_rate: 16000 },
      output: { encoding: "linear16", sample_rate: 24000, container: "none" },
    },
    agent: {
      listen: { provider: { type: "deepgram", model: "flux-general-en", version: "v2" } },
      think: { provider: { type: "open_ai", model: "gpt-4o-mini" }, // a "Standard" tier model
               prompt: "You are a helpful assistant. Keep answers short." },
      speak: { provider: { type: "deepgram", model: "aura-2-thalia-en" } },
      greeting: "Hi! How can I help?",
    },
  }));
  setInterval(() => ws.send(JSON.stringify({ type: "KeepAlive" })), 5000);
});

// Send raw linear16 mic audio as binary frames
export function sendPcm(buf) { ws.send(buf); }

ws.on("message", (data, isBinary) => {
  if (isBinary) return play(data);          // agent speech (raw PCM)
  const ev = JSON.parse(data.toString());
  if (ev.type === "UserStartedSpeaking") stopPlayback(); // barge-in
  if (ev.type === "Error") console.error(ev);
});
function play(pcm) {} function stopPlayback() {}

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Model choice changes the tier

Picking an Advanced-tier LLM (e.g. GPT-5, Claude Sonnet) moves the session to $0.163/min, more than double Standard. The tier follows the model, so a config change can double cost.

Connection time is billed

The meter runs while the WebSocket is open, including silence and hold. Close sockets promptly and avoid idle KeepAlive loops.

BYO tiers are not all-in

BYO LLM/TTS rates exclude what you pay OpenAI, Bedrock, Groq, ElevenLabs etc. Bedrock and Groq always require your own endpoint and credentials.

Promotional rates expired

Third-party pages still quote promo or older rates (e.g. $0.08 Standard, $0.056 BYO LLM); some promos ended 2026-09-12. Use deepgram.com/pricing.

Concurrency ceiling

45 concurrent sessions on PAYG and 60 on Growth; larger contact centres need Enterprise.

Deprecated fields

agent.language is deprecated; set language on listen.provider and speak.provider. Some listed models are already marked deprecated.

Cascaded, not native speech-to-speech

STT, LLM and TTS are separate stages, so paralinguistic cues (tone, laughter) are not passed to the LLM.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Concurrency: up to 45 WebSocket agent sessions on Pay As You Go, up to 60 on Growth
  • Billing runs for the whole WebSocket connection

Models and products

NameStatusNotes
STT: Deepgram Flux / Nova (e.g. flux-general-en, flux-general-multi)GA
LLM Standard tierGAe.g. gpt-5-mini, gpt-5.4-mini, gpt-4o-mini, claude-haiku-4-5, gemini-3.5-flash, gemini-2.5-flash, nemotron-3-nano, groq openai/gpt-oss-20b (BYO).
LLM Advanced tierGAe.g. gpt-5.5, gpt-5.4, gpt-5, gpt-4.1, gpt-4o, claude-sonnet-5, claude-sonnet-4-6, gemini-3-pro-preview.
TTS: Deepgram Aura-2 / Flux TTS, or eleven_labs, cartesia, open_ai, aws_pollyGANon-Deepgram TTS requires your own endpoint/key (BYO TTS tiers).

Docs and sources

Docs

Sources used

Not fully verified

Latency figures; compliance certifications; Growth plan entry requirements; whether 'Token' vs 'Bearer' is required for every key type.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?