GA Cartesia

Cartesia Sonic

State-space-model TTS built for voice agents; one WebSocket carries many contexts with continuation, word and phoneme timestamps, and native mu-law/a-law output. Current model is Sonic 3.6.

Est. per minute$0.034 - 0.059

Overview

Best for: Latency-critical voice agents and telephony (native mu-law 8 kHz), with timestamps for interruption handling.

At a glance

$/1M chars$50
Free tierYes
Free tier commercialNo
CloningYes
Instant cloneYes
Text stream inYes
TimestampsYes
EmotionYes
8 kHz phoneYes
Latency ms90
Languages44
Concurrency3
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

Price is the Pro plan effective rate ($5 for 100K credits, 1 credit per char); Pro overage $65/1M, Scale about $37/1M. Latency ~90 ms is a vendor claim quoted by third parties. Concurrency 3 on Pro (Free 2, Scale 15). Break tags supported but not full SSML. Free plan not listed for commercial use.

Audio in

Text (no SSML required; break tags supported)

Audio out

Raw container; pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw at 8000/16000/22050/24000/44100/48000 Hz on the WebSocket

Languages

44 (Sonic 3.6)

Voices

Voice library plus instant voice cloning (Pro and up) and professional cloning (Startup and up)

Latency

Third-party sources quote Cartesia's claims of ~90 ms time-to-first-audio for Sonic 3 and ~40 ms for Sonic Turbo; Sonic 3.6 described as sub-90 ms. Independent benchmarks reported 128-166 ms medians including network for Sonic 3/3.5.

Regions

Not specified on the pages checked

Compliance

Not verified in this pass (enterprise custom terms available).

Features

  • input streaming via contexts (continue flag)
  • word timestamps
  • phoneme timestamps
  • emotion, speed and volume controls (generation_config)
  • voice cloning
  • pronunciation dictionaries
  • cancel per context
  • access tokens for browser clients

Pricing

WhatPriceUnitNotes
Free$0monthly20K credits; TTS concurrency 2; no cloning
Pro$5/monthmonthly100K credits; TTS concurrency 3; instant cloning; overage $65 per 1M credits
Startup$49/monthmonthly1.25M credits; TTS concurrency 5; overage $45 per 1M credits
Scale$299/monthmonthly8M credits; TTS concurrency 15; overage $38 per 1M credits
TTS credit rate1 creditper character~750-800 credits per minute of audio (vendor FAQ); each break tag = 1 credit
How the per-minute estimate was worked out

900 chars/min at 1 credit/char. Low = Scale plan effective ~$37.4/1M; high = Pro overage $65/1M.

Free tier: 20K credits/month, overages blocked; commercial use not listed for Free (Pro and up: yes).

Source: cartesia.ai

Setup

  1. Create an API key in the Cartesia playground/console.
  2. Pick a voice ID.
  3. Open wss://api.cartesia.ai/tts/websocket with X-API-Key and a cartesia_version.
  4. Send generation requests sharing one context_id; set continue:true on every chunk except the last.
  5. Read base64 'chunk' messages until 'done'.

Endpoint

wss://api.cartesia.ai/tts/websocket

Authentication

X-API-Key header (server) or access_token query param (browser); cartesia_version required (example 2026-08-14)

Quick start python

# pip install websockets   (streams text chunks into one context)
import asyncio, base64, json, os, websockets

URL = "wss://api.cartesia.ai/tts/websocket?cartesia_version=2026-08-14"
HDR = {"X-API-Key": os.environ["CARTESIA_API_KEY"]}

async def main():
    async with websockets.connect(URL, additional_headers=HDR) as ws:
        chunks = ["Hello there. ", "This arrives ", "token by token."]
        for i, text in enumerate(chunks):
            await ws.send(json.dumps({
                "model_id": "sonic-3.6",
                "transcript": text,
                "voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
                "output_format": {"container": "raw", "encoding": "pcm_s16le", "sample_rate": 24000},
                "context_id": "turn-1",
                "continue": i < len(chunks) - 1,
            }))
        with open("out_24k_s16le.pcm", "wb") as f:
            async for raw in ws:
                msg = json.loads(raw)
                if msg.get("type") == "chunk":
                    f.write(base64.b64decode(msg["data"]))
                elif msg.get("type") in ("done", "error"):
                    break

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Pin a dated snapshot

sonic-3.6 is an alias that moves to the newest stable snapshot. Pin sonic-3.6-2026-08-27 (or the current dated ID) for production so voice behaviour does not change silently.

Low TTS concurrency on small plans

Free allows 2 and Pro 3 concurrent TTS generations; a handful of simultaneous calls will hit 429s. Size the plan to peak simultaneous speakers, not monthly volume.

Last chunk must set continue:false

If you never send a final request with continue:false (or flush), the server waits up to max_buffer_delay_ms (default 3 s) before speaking the tail of the turn.

Raw container only on the WebSocket

The socket returns headerless raw audio; you must know the encoding and sample rate to play or wrap it as WAV.

Break tags cost credits

Each break tag counts as 1 credit; markup-heavy text costs slightly more than the spoken characters.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • TTS concurrency: Free 2, Pro 3, Startup 5, Scale 15
  • max_buffer_delay_ms default 3000 (0-5000)
  • voice and output_format must stay constant within a context
  • Overages must be enabled; Free stops at the credit cap

Models and products

NameStatusNotes
sonic-3.6GA (alias to latest stable snapshot)44 languages; backwards compatible with Sonic 3.5.
sonic-3.6-2026-08-27Stable snapshotPin this in production so behaviour does not change under you.
sonic-3.5, sonic-3Older, still accepted on the WebSocket
sonic-latest / sonic-previewBetasonic-preview may change without notice; not for production.
sonic-2, sonic-turbo, sonicOlder modelsListed only as older models; not in the current WebSocket model_id list.

Docs and sources

Docs

Sources used

Not fully verified

Latency numbers (from third-party write-ups of Cartesia claims); whether cartesia_version must be header or query (snippet uses query); regions; voice count.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?