Deepgram Nova-3 and Flux (streaming)
Low-cost WebSocket streaming STT. Nova-3 is the general model for captions, meetings and multilingual audio; Flux is a separate voice-agent model with built-in end-of-turn detection on its own /v2 endpoint.
Overview
Best for: Cheap, high-concurrency English or multilingual live captions (Nova-3) and voice agents that want turn detection from the STT itself (Flux).
At a glance
Price is Nova-3 monolingual PAYG ($0.0048/min, marked promotional); Flux English is $0.39/hr. Latency is Flux end-of-turn (~260 ms); no Nova-3 figure. 59 languages for Nova-3; code-switching across 10. Diarization, keyterms and redaction are paid add-ons. HIPAA, SOC 2 and EU endpoint are vendor claims, not re-verified. 10 s idle timeout, no session cap stated.
Raw linear16, linear32, mulaw, alaw, opus, ogg-opus need encoding + sample_rate (8000, 16000, 24000, 44100, 48000 listed for Flux); containerized WAV/Ogg/WebM can be sent without those params. Flux strongly recommends 80 ms chunks (2560 bytes at 16 kHz linear16).
n/a
Nova-3: multilingual code-switching mode for 10 languages plus a long list of single-language codes (see Models and Languages page). Flux: English, or 10 languages with flux-general-multi.
Vendor claim: Flux end-of-turn detection around 260 ms. No published latency figure for Nova-3 found in this pass.
Hosted API default region; EU endpoint and self-hosted deployment exist (not re-verified in this pass).
Vendor markets HIPAA (BAA) and SOC 2 and offers self-hosted/on-prem deployment; not re-verified in this pass, confirm with sales.
Features
- interim results (Nova-3)
- endpointing and utterance end (Nova-3)
- model-integrated end-of-turn with EagerEndOfTurn / TurnResumed (Flux)
- diarization (paid add-on in streaming, Nova-3)
- keyterm prompting (paid add-on)
- redaction (paid add-on)
- smart formatting (included)
- word timestamps
- Configure message to change Flux settings mid-stream
- ForceEndTurn for bring-your-own turn detection
Pricing
| What | Price | Unit |
|---|---|---|
| Nova-3 monolingual streaming | $0.0048 / $0.0042 | per minute (Pay As You Go / Growth) |
| Nova-3 multilingual streaming | $0.0058 / $0.0050 | per minute (PAYG / Growth) |
| Flux English streaming | $0.0065 / $0.0057 | per minute (PAYG / Growth) |
| Flux Multilingual streaming | $0.0078 / $0.0068 | per minute (PAYG / Growth) |
| Speaker diarization (streaming) | +$0.0020 / +$0.0017 | per minute |
| Keyterm prompting | +$0.0013 / +$0.0012 | per minute |
| Redaction | +$0.0020 / +$0.0017 | per minute |
| Entity detection | +$0.0017 | per minute |
Low = Nova-3 mono on Growth; high = Flux Multilingual PAYG ($0.0078) + diarization ($0.0020) where supported. Add-ons stack.
Free tier: $200 free credit for new accounts, does not expire until used (pricing page).
Source: deepgram.com
Setup
- Create an API key in the Deepgram console.
- Open wss://api.deepgram.com/v1/listen with model=nova-3 (or wss://api.deepgram.com/v2/listen?model=flux-general-en for Flux).
- Send header Authorization: Token <key>; for browsers mint a short-lived JWT server-side (token-based auth, 30 s TTL).
- Stream raw audio frames, send KeepAlive during silence, send {"type":"CloseStream"} to flush and close.
Endpoint
wss://api.deepgram.com/v1/listen (Nova-3) ; wss://api.deepgram.com/v2/listen (Flux)
Authentication
Authorization: Token <API_KEY> header. Browser: temporary JWT from the token-based auth endpoint (30 second TTL to open the socket).
Quick start python
import asyncio, json, os, websockets # pip install websockets>=14
URL = ("wss://api.deepgram.com/v1/listen?model=nova-3&encoding=linear16"
"&sample_rate=16000&interim_results=true&endpointing=300")
async def main():
hdr = {"Authorization": f"Token {os.environ['DEEPGRAM_API_KEY']}"}
async with websockets.connect(URL, additional_headers=hdr) as ws:
async def send():
with open("audio_16k_mono.raw", "rb") as f: # 16-bit PCM mono
while chunk := f.read(3200): # 100 ms
await ws.send(chunk)
await asyncio.sleep(0.1) # pace at real time
await ws.send(json.dumps({"type": "CloseStream"}))
asyncio.create_task(send())
async for msg in ws:
d = json.loads(msg)
if d.get("type") == "Results":
text = d["channel"]["alternatives"][0]["transcript"]
if text:
print("FINAL" if d["is_final"] else "partial", text)
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Flux needs /v2, not /v1
Flux only works on wss://api.deepgram.com/v2/listen with model=flux-general-en or flux-general-multi. model=flux is invalid and the v1 endpoint will not accept it. Its message types (TurnInfo, EndOfTurn) differ from Nova-3 Results messages, so it is not a drop-in swap.
10-second idle timeout
If neither audio nor a KeepAlive text frame arrives for 10 s the socket closes with NET-0001. Send KeepAlive every 3-5 s during mute or hold; send it as a text frame, not binary.
Diarization is extra on streaming only
Speaker diarization costs +$0.0020/min PAYG on streaming but is included on pre-recorded. Flux's feature matrix does not list diarization at all.
Streaming prices are marked promotional
The pricing page shows the streaming rates with strikethrough figures labelled limited-time promotional. Budget for the possibility that the struck-through (higher) rates return.
Eager end-of-turn multiplies LLM calls
Setting eager_eot_threshold enables EagerEndOfTurn/TurnResumed; Deepgram's own docs warn this can raise LLM calls by 50-70% because some speculative turns get cancelled.
Add-ons stack per minute
Keyterm prompting, redaction and entity detection each add their own per-minute charge on top of the base model rate.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Concurrency: 150 WebSocket STT streams on Pay As You Go, 225 on Growth (pricing page).
- Connection closes with NET-0001 if no audio or KeepAlive for 10 seconds; send KeepAlive (text frame) every 3-5 s.
- Whisper Cloud limited to 5 concurrent requests.
Models and products
| Name | Status |
|---|---|
| nova-3 (nova-3-general) | GA |
| nova-3-medical | GA |
| nova-3-pharma | Listed in docs |
| flux-general-en | GA |
| flux-general-multi | Listed in docs |
| nova-2 / enhanced / base | Legacy |
Docs and sources
Docs
Sources used
- deepgram.com/pricing
- developers.deepgram.com/docs/flux/quickstart
- developers.deepgram.com/docs/flux/feature-overview.md
- developers.deepgram.com/docs/audio-keep-alive.md
- developers.deepgram.com/docs/models-languages-overview.md
Compliance claims, regional endpoints and Nova-3 latency numbers were not re-checked. Whether streaming billing counts wall-clock connection time or audio sent was not confirmed on a primary page.