Cartesia Sonic
State-space-model TTS built for voice agents; one WebSocket carries many contexts with continuation, word and phoneme timestamps, and native mu-law/a-law output. Current model is Sonic 3.6.
Overview
Best for: Latency-critical voice agents and telephony (native mu-law 8 kHz), with timestamps for interruption handling.
At a glance
Price is the Pro plan effective rate ($5 for 100K credits, 1 credit per char); Pro overage $65/1M, Scale about $37/1M. Latency ~90 ms is a vendor claim quoted by third parties. Concurrency 3 on Pro (Free 2, Scale 15). Break tags supported but not full SSML. Free plan not listed for commercial use.
Text (no SSML required; break tags supported)
Raw container; pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw at 8000/16000/22050/24000/44100/48000 Hz on the WebSocket
44 (Sonic 3.6)
Voice library plus instant voice cloning (Pro and up) and professional cloning (Startup and up)
Third-party sources quote Cartesia's claims of ~90 ms time-to-first-audio for Sonic 3 and ~40 ms for Sonic Turbo; Sonic 3.6 described as sub-90 ms. Independent benchmarks reported 128-166 ms medians including network for Sonic 3/3.5.
Not specified on the pages checked
Not verified in this pass (enterprise custom terms available).
Features
- input streaming via contexts (continue flag)
- word timestamps
- phoneme timestamps
- emotion, speed and volume controls (generation_config)
- voice cloning
- pronunciation dictionaries
- cancel per context
- access tokens for browser clients
Pricing
| What | Price | Unit |
|---|---|---|
| Free | $0 | monthly |
| Pro | $5/month | monthly |
| Startup | $49/month | monthly |
| Scale | $299/month | monthly |
| TTS credit rate | 1 credit | per character |
900 chars/min at 1 credit/char. Low = Scale plan effective ~$37.4/1M; high = Pro overage $65/1M.
Free tier: 20K credits/month, overages blocked; commercial use not listed for Free (Pro and up: yes).
Source: cartesia.ai
Setup
- Create an API key in the Cartesia playground/console.
- Pick a voice ID.
- Open wss://api.cartesia.ai/tts/websocket with X-API-Key and a cartesia_version.
- Send generation requests sharing one context_id; set continue:true on every chunk except the last.
- Read base64 'chunk' messages until 'done'.
Endpoint
wss://api.cartesia.ai/tts/websocket
Authentication
X-API-Key header (server) or access_token query param (browser); cartesia_version required (example 2026-08-14)
Quick start python
# pip install websockets (streams text chunks into one context)
import asyncio, base64, json, os, websockets
URL = "wss://api.cartesia.ai/tts/websocket?cartesia_version=2026-08-14"
HDR = {"X-API-Key": os.environ["CARTESIA_API_KEY"]}
async def main():
async with websockets.connect(URL, additional_headers=HDR) as ws:
chunks = ["Hello there. ", "This arrives ", "token by token."]
for i, text in enumerate(chunks):
await ws.send(json.dumps({
"model_id": "sonic-3.6",
"transcript": text,
"voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
"output_format": {"container": "raw", "encoding": "pcm_s16le", "sample_rate": 24000},
"context_id": "turn-1",
"continue": i < len(chunks) - 1,
}))
with open("out_24k_s16le.pcm", "wb") as f:
async for raw in ws:
msg = json.loads(raw)
if msg.get("type") == "chunk":
f.write(base64.b64decode(msg["data"]))
elif msg.get("type") in ("done", "error"):
break
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Pin a dated snapshot
sonic-3.6 is an alias that moves to the newest stable snapshot. Pin sonic-3.6-2026-08-27 (or the current dated ID) for production so voice behaviour does not change silently.
Low TTS concurrency on small plans
Free allows 2 and Pro 3 concurrent TTS generations; a handful of simultaneous calls will hit 429s. Size the plan to peak simultaneous speakers, not monthly volume.
Last chunk must set continue:false
If you never send a final request with continue:false (or flush), the server waits up to max_buffer_delay_ms (default 3 s) before speaking the tail of the turn.
Raw container only on the WebSocket
The socket returns headerless raw audio; you must know the encoding and sample rate to play or wrap it as WAV.
Break tags cost credits
Each break tag counts as 1 credit; markup-heavy text costs slightly more than the spoken characters.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- TTS concurrency: Free 2, Pro 3, Startup 5, Scale 15
- max_buffer_delay_ms default 3000 (0-5000)
- voice and output_format must stay constant within a context
- Overages must be enabled; Free stops at the credit cap
Models and products
| Name | Status |
|---|---|
| sonic-3.6 | GA (alias to latest stable snapshot) |
| sonic-3.6-2026-08-27 | Stable snapshot |
| sonic-3.5, sonic-3 | Older, still accepted on the WebSocket |
| sonic-latest / sonic-preview | Beta |
| sonic-2, sonic-turbo, sonic | Older models |
Docs and sources
Docs
Sources used
- cartesia.ai/pricing
- docs.cartesia.ai/build-with-cartesia/tts-models/latest
- docs.cartesia.ai/api-reference/tts/websocket
- humannessindex.vapi.ai/models/cartesia-sonic-3-5
Latency numbers (from third-party write-ups of Cartesia claims); whether cartesia_version must be header or query (snippet uses query); regions; voice count.