GA Rime Labs

Rime TTS (Coda, Mist v3)

Conversational TTS aimed at contact-centre agents, with a JSON WebSocket (/ws3), word timestamps, native mu-law and an on-prem option. Arcana was retired from the cloud on 2026-08-15 and now routes to Coda.

Est. per minute$0.027 - 0.045
1 high-severity warning

Overview

Best for: US contact-centre and telephony agents that need natural conversational English, mu-law output and an on-prem path.

At a glance

$/1M chars$50
Free tierYes
Text stream inYes
TimestampsYes
8 kHz phoneYes
Latency ms200
Languages8
Concurrency20
WebRTCNo
WebSocketYes
gRPCNo
HIPAAYes
SOC 2Yes
Self-hostYes
Open weightsNo

Figures are Coda ($50/1M, 8 languages); Mist v3 is $30/1M with 4 languages and TTFB well below 100 ms. Cloud claim sub-200 ms end to end. Concurrency 20 on Starter. BAA and SOC 2 reports on Enterprise. Free allowance unclear (800 vs 3,000 minutes).

Audio in

Text (custom pronunciation via spell() and dictionaries)

Audio out

pcm, mp3, mulaw (audioFormat); sample rates per model reference

Languages

Coda: 8; Mist v3: 4

Voices

Catalog voices; custom voice clones on Enterprise

Latency

Vendor: sub-200 ms end-to-end via cloud API; Mist v3 typical TTFB well below 100 ms; Coda sub-100 ms model latency on GPU engine (self-hosted) plus 25-50 ms network from most of the US.

Regions

users-ws.rime.ai plus users-east-ws.rime.ai (us-east-1) and users-west regional hosts

Compliance

Enterprise: BAA and SOC 2 reports; optional request-data retention controls (disabled by default).

Features

  • input streaming (/ws3)
  • word timestamps
  • contextId echo
  • flush / clear / eos operations
  • segment modes (never, bySentence, immediate)
  • on-prem Docker/Kubernetes
  • MCP server

Pricing

WhatPriceUnitNotes
Mist v3$0.03per 1K charactersStarter; vendor says ~$0.03/minute
Coda$0.05per 1K charactersStarter; ~$0.05/minute
EnterpriseCustomUnlimited concurrency, BAA, SOC 2 reports, on-prem/VPC
How the per-minute estimate was worked out

900 chars/min; Mist v3 vs Coda at Starter rates

Free tier: Page conflicts: '~800 minutes free (about 800k characters)' vs FAQ '3,000 free minutes'

Source: rime.ai

Setup

  1. Create an API key at rime.ai.
  2. Connect to wss://users-ws.rime.ai/ws3 with speaker, modelId and audioFormat query params and an Authorization: Bearer header.
  3. Send {text} messages, then {operation:'flush'} or {operation:'eos'}.
  4. Decode base64 'chunk' events; use 'timestamps' for barge-in.

Endpoint

wss://users-ws.rime.ai/ws3

Authentication

Authorization: Bearer <RIME_API_KEY> header (browsers cannot set it; proxy via backend)

Quick start javascript

// npm i ws
import WebSocket from "ws";
import fs from "fs";

const url = "wss://users-ws.rime.ai/ws3?speaker=YOUR_SPEAKER&modelId=mistv3&audioFormat=pcm";
const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.RIME_API_KEY}` } });
const out = fs.createWriteStream("out.pcm");

ws.on("open", () => {
  ws.send(JSON.stringify({ text: "Hello there. ", contextId: "turn-1" }));
  ws.send(JSON.stringify({ text: "Thanks for calling today." }));
  ws.send(JSON.stringify({ operation: "eos" }));   // speak the rest, then close
});
ws.on("message", (raw) => {
  const m = JSON.parse(raw.toString());
  if (m.type === "chunk") out.write(Buffer.from(m.data, "base64"));
  if (m.type === "error") console.error(m);
});
ws.on("close", () => out.end());

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Arcana is gone from the cloud

Since 2026-08-15 arcana, arcanav2 and arcanav3 requests are silently routed to Coda. Voices and pricing may differ from what you tested; re-evaluate.

Endpoint/model matrix is inconsistent in docs

One reference lists mistv2 as not supported on /ws3 while the overview lists it. Test your exact modelId on /ws3 before shipping.

Free allowance is unclear

The pricing page says ~800 minutes free while the FAQ says 3,000 minutes. Do not plan around either number.

Single context per socket

Rime does not track multiple simultaneous contextIds; open one socket per concurrent speaker.

Per-minute price assumes ~1K chars/min

Rime equates $0.03/1K chars with ~$0.03/min; dense text costs more.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Starter: 20 concurrent TTS generations
  • Rime keeps only one active contextId at a time per socket

Models and products

NameStatusNotes
codaGA (launched 2026-05-19)Successor to Arcana; English, French, German, Japanese, Portuguese, Spanish, plus Arabic and Hindi (2026-08-04).
mistv3GA (2026-04)Lowest cloud latency; English, French, German, Spanish.
mistv2 / mistv1LegacyStill served; /ws2 limited to these.
arcana / arcanav2 / arcanav3Retired from cloud 2026-08-15Requests now route to Coda.

Docs and sources

Docs

Sources used

Not fully verified

Per-model sample rates; speaker names for mistv3 (placeholder in snippet).

Similar text-to-speech APIs

Spotted a wrong price or a dead link?