Rime TTS (Coda, Mist v3)
Conversational TTS aimed at contact-centre agents, with a JSON WebSocket (/ws3), word timestamps, native mu-law and an on-prem option. Arcana was retired from the cloud on 2026-08-15 and now routes to Coda.
Overview
Best for: US contact-centre and telephony agents that need natural conversational English, mu-law output and an on-prem path.
At a glance
Figures are Coda ($50/1M, 8 languages); Mist v3 is $30/1M with 4 languages and TTFB well below 100 ms. Cloud claim sub-200 ms end to end. Concurrency 20 on Starter. BAA and SOC 2 reports on Enterprise. Free allowance unclear (800 vs 3,000 minutes).
Text (custom pronunciation via spell() and dictionaries)
pcm, mp3, mulaw (audioFormat); sample rates per model reference
Coda: 8; Mist v3: 4
Catalog voices; custom voice clones on Enterprise
Vendor: sub-200 ms end-to-end via cloud API; Mist v3 typical TTFB well below 100 ms; Coda sub-100 ms model latency on GPU engine (self-hosted) plus 25-50 ms network from most of the US.
users-ws.rime.ai plus users-east-ws.rime.ai (us-east-1) and users-west regional hosts
Enterprise: BAA and SOC 2 reports; optional request-data retention controls (disabled by default).
Features
- input streaming (/ws3)
- word timestamps
- contextId echo
- flush / clear / eos operations
- segment modes (never, bySentence, immediate)
- on-prem Docker/Kubernetes
- MCP server
Pricing
| What | Price | Unit |
|---|---|---|
| Mist v3 | $0.03 | per 1K characters |
| Coda | $0.05 | per 1K characters |
| Enterprise | Custom |
900 chars/min; Mist v3 vs Coda at Starter rates
Free tier: Page conflicts: '~800 minutes free (about 800k characters)' vs FAQ '3,000 free minutes'
Source: rime.ai
Setup
- Create an API key at rime.ai.
- Connect to wss://users-ws.rime.ai/ws3 with speaker, modelId and audioFormat query params and an Authorization: Bearer header.
- Send {text} messages, then {operation:'flush'} or {operation:'eos'}.
- Decode base64 'chunk' events; use 'timestamps' for barge-in.
Endpoint
wss://users-ws.rime.ai/ws3
Authentication
Authorization: Bearer <RIME_API_KEY> header (browsers cannot set it; proxy via backend)
Quick start javascript
// npm i ws
import WebSocket from "ws";
import fs from "fs";
const url = "wss://users-ws.rime.ai/ws3?speaker=YOUR_SPEAKER&modelId=mistv3&audioFormat=pcm";
const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.RIME_API_KEY}` } });
const out = fs.createWriteStream("out.pcm");
ws.on("open", () => {
ws.send(JSON.stringify({ text: "Hello there. ", contextId: "turn-1" }));
ws.send(JSON.stringify({ text: "Thanks for calling today." }));
ws.send(JSON.stringify({ operation: "eos" })); // speak the rest, then close
});
ws.on("message", (raw) => {
const m = JSON.parse(raw.toString());
if (m.type === "chunk") out.write(Buffer.from(m.data, "base64"));
if (m.type === "error") console.error(m);
});
ws.on("close", () => out.end());
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Arcana is gone from the cloud
Since 2026-08-15 arcana, arcanav2 and arcanav3 requests are silently routed to Coda. Voices and pricing may differ from what you tested; re-evaluate.
Endpoint/model matrix is inconsistent in docs
One reference lists mistv2 as not supported on /ws3 while the overview lists it. Test your exact modelId on /ws3 before shipping.
Free allowance is unclear
The pricing page says ~800 minutes free while the FAQ says 3,000 minutes. Do not plan around either number.
Single context per socket
Rime does not track multiple simultaneous contextIds; open one socket per concurrent speaker.
Per-minute price assumes ~1K chars/min
Rime equates $0.03/1K chars with ~$0.03/min; dense text costs more.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Starter: 20 concurrent TTS generations
- Rime keeps only one active contextId at a time per socket
Models and products
| Name | Status |
|---|---|
| coda | GA (launched 2026-05-19) |
| mistv3 | GA (2026-04) |
| mistv2 / mistv1 | Legacy |
| arcana / arcanav2 / arcanav3 | Retired from cloud 2026-08-15 |
Docs and sources
Docs
Sources used
Per-model sample rates; speaker names for mistv3 (placeholder in snippet).