GA Deepgram

Deepgram Nova-3 and Flux (streaming)

Low-cost WebSocket streaming STT. Nova-3 is the general model for captions, meetings and multilingual audio; Flux is a separate voice-agent model with built-in end-of-turn detection on its own /v2 endpoint.

Est. per minute$0.0042 - 0.0098
2 high-severity warnings

Overview

Best for: Cheap, high-concurrency English or multilingual live captions (Nova-3) and voice agents that want turn detection from the STT itself (Flux).

At a glance

$/hour$0.288
Free tierYes
Free credit $$200
Live speakersYes
Turn detectYes
KeytermsYes
PII redactYes
PartialsYes
Mixed langsYes
8 kHz phoneYes
Latency ms260
Languages59
Concurrency150
WebRTCNo
WebSocketYes
gRPCNo
HIPAAYes
SOC 2Yes
EU dataYes
Self-hostYes
Open weightsNo

Price is Nova-3 monolingual PAYG ($0.0048/min, marked promotional); Flux English is $0.39/hr. Latency is Flux end-of-turn (~260 ms); no Nova-3 figure. 59 languages for Nova-3; code-switching across 10. Diarization, keyterms and redaction are paid add-ons. HIPAA, SOC 2 and EU endpoint are vendor claims, not re-verified. 10 s idle timeout, no session cap stated.

Audio in

Raw linear16, linear32, mulaw, alaw, opus, ogg-opus need encoding + sample_rate (8000, 16000, 24000, 44100, 48000 listed for Flux); containerized WAV/Ogg/WebM can be sent without those params. Flux strongly recommends 80 ms chunks (2560 bytes at 16 kHz linear16).

Audio out

n/a

Languages

Nova-3: multilingual code-switching mode for 10 languages plus a long list of single-language codes (see Models and Languages page). Flux: English, or 10 languages with flux-general-multi.

Latency

Vendor claim: Flux end-of-turn detection around 260 ms. No published latency figure for Nova-3 found in this pass.

Regions

Hosted API default region; EU endpoint and self-hosted deployment exist (not re-verified in this pass).

Compliance

Vendor markets HIPAA (BAA) and SOC 2 and offers self-hosted/on-prem deployment; not re-verified in this pass, confirm with sales.

Features

  • interim results (Nova-3)
  • endpointing and utterance end (Nova-3)
  • model-integrated end-of-turn with EagerEndOfTurn / TurnResumed (Flux)
  • diarization (paid add-on in streaming, Nova-3)
  • keyterm prompting (paid add-on)
  • redaction (paid add-on)
  • smart formatting (included)
  • word timestamps
  • Configure message to change Flux settings mid-stream
  • ForceEndTurn for bring-your-own turn detection

Pricing

WhatPriceUnitNotes
Nova-3 monolingual streaming$0.0048 / $0.0042per minute (Pay As You Go / Growth)Pre-recorded is cheaper: $0.0043 / $0.0036.
Nova-3 multilingual streaming$0.0058 / $0.0050per minute (PAYG / Growth)Pre-recorded $0.0052 / $0.0043.
Flux English streaming$0.0065 / $0.0057per minute (PAYG / Growth)
Flux Multilingual streaming$0.0078 / $0.0068per minute (PAYG / Growth)
Speaker diarization (streaming)+$0.0020 / +$0.0017per minuteIncluded free on pre-recorded, charged on streaming.
Keyterm prompting+$0.0013 / +$0.0012per minute
Redaction+$0.0020 / +$0.0017per minute
Entity detection+$0.0017per minute
How the per-minute estimate was worked out

Low = Nova-3 mono on Growth; high = Flux Multilingual PAYG ($0.0078) + diarization ($0.0020) where supported. Add-ons stack.

Free tier: $200 free credit for new accounts, does not expire until used (pricing page).

Source: deepgram.com

Setup

  1. Create an API key in the Deepgram console.
  2. Open wss://api.deepgram.com/v1/listen with model=nova-3 (or wss://api.deepgram.com/v2/listen?model=flux-general-en for Flux).
  3. Send header Authorization: Token <key>; for browsers mint a short-lived JWT server-side (token-based auth, 30 s TTL).
  4. Stream raw audio frames, send KeepAlive during silence, send {"type":"CloseStream"} to flush and close.

Endpoint

wss://api.deepgram.com/v1/listen (Nova-3) ; wss://api.deepgram.com/v2/listen (Flux)

Authentication

Authorization: Token <API_KEY> header. Browser: temporary JWT from the token-based auth endpoint (30 second TTL to open the socket).

Quick start python

import asyncio, json, os, websockets  # pip install websockets>=14

URL = ("wss://api.deepgram.com/v1/listen?model=nova-3&encoding=linear16"
       "&sample_rate=16000&interim_results=true&endpointing=300")

async def main():
    hdr = {"Authorization": f"Token {os.environ['DEEPGRAM_API_KEY']}"}
    async with websockets.connect(URL, additional_headers=hdr) as ws:
        async def send():
            with open("audio_16k_mono.raw", "rb") as f:  # 16-bit PCM mono
                while chunk := f.read(3200):  # 100 ms
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)  # pace at real time
            await ws.send(json.dumps({"type": "CloseStream"}))
        asyncio.create_task(send())
        async for msg in ws:
            d = json.loads(msg)
            if d.get("type") == "Results":
                text = d["channel"]["alternatives"][0]["transcript"]
                if text:
                    print("FINAL" if d["is_final"] else "partial", text)

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Flux needs /v2, not /v1

Flux only works on wss://api.deepgram.com/v2/listen with model=flux-general-en or flux-general-multi. model=flux is invalid and the v1 endpoint will not accept it. Its message types (TurnInfo, EndOfTurn) differ from Nova-3 Results messages, so it is not a drop-in swap.

10-second idle timeout

If neither audio nor a KeepAlive text frame arrives for 10 s the socket closes with NET-0001. Send KeepAlive every 3-5 s during mute or hold; send it as a text frame, not binary.

Diarization is extra on streaming only

Speaker diarization costs +$0.0020/min PAYG on streaming but is included on pre-recorded. Flux's feature matrix does not list diarization at all.

Streaming prices are marked promotional

The pricing page shows the streaming rates with strikethrough figures labelled limited-time promotional. Budget for the possibility that the struck-through (higher) rates return.

Eager end-of-turn multiplies LLM calls

Setting eager_eot_threshold enables EagerEndOfTurn/TurnResumed; Deepgram's own docs warn this can raise LLM calls by 50-70% because some speculative turns get cancelled.

Add-ons stack per minute

Keyterm prompting, redaction and entity detection each add their own per-minute charge on top of the base model rate.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Concurrency: 150 WebSocket STT streams on Pay As You Go, 225 on Growth (pricing page).
  • Connection closes with NET-0001 if no audio or KeepAlive for 10 seconds; send KeepAlive (text frame) every 3-5 s.
  • Whisper Cloud limited to 5 concurrent requests.

Models and products

NameStatusNotes
nova-3 (nova-3-general)GAGeneral streaming and batch model. language=multi gives code-switching across 10 languages (en, es, fr, de, hi, ru, pt, ja, it, nl); many single-language codes also listed.
nova-3-medicalGAMedical vocabulary variant; English plus the same 10-language multi mode.
nova-3-pharmaListed in docsEnglish only, pharma vocabulary.
flux-general-enGAConversational STT for voice agents with model-integrated end-of-turn. Only on wss://api.deepgram.com/v2/listen (v1 does not work).
flux-general-multiListed in docsFlux multilingual, 10 languages, language_hint supported.
nova-2 / enhanced / baseLegacyStill available at old rates (Nova-2 streaming $0.35/hr per pricing FAQ). Use Nova-2 only for languages Nova-3 lacks.

Docs and sources

Docs

Sources used

Not fully verified

Compliance claims, regional endpoints and Nova-3 latency numbers were not re-checked. Whether streaming billing counts wall-clock connection time or audio sent was not confirmed on a primary page.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?