GA Resemble AI

Resemble AI TTS (Chatterbox models)

Voice cloning company whose hosted API serves its Chatterbox models (e.g. chatterbox-turbo). Its WebSocket streams audio out but takes a whole text/SSML payload per request. The public pricing page now covers deepfake detection only.

Est. per minuten/a
2 high-severity warnings

Overview

Best for: Teams that want hosted Chatterbox with cloning plus deepfake detection/watermarking from one vendor.

At a glance

CloningYes
Text stream inNo
TimestampsYes
SSMLYes
Latency ms200
Languages1
Concurrency20
WebRTCNo
WebSocketYes
gRPCNo
SOC 2Yes
Self-hostYes
Open weightsYes

Chatterbox Turbo (English) sub-200 ms vendor claim; Chatterbox Multilingual open weights cover 23+ languages. TTS pricing not published. WebSocket needs the Business plan; 20 sessions is the default there. Timestamps are character and phoneme level. SOC 2 documentation mentioned for enterprise.

Audio in

Text or SSML, max 3,000 characters excluding tags

Audio out

wav/pcm etc.; example uses wav 32 kHz PCM_32; JSON (base64) or binary frames

Languages

English for Turbo; Chatterbox Multilingual covers 23+ languages

Voices

Voice cloning (rapid and professional)

Latency

Vendor (Chatterbox README): hosted service sub-200 ms.

Regions

Not specified

Compliance

Enterprise on-prem available; SOC 2 documentation mentioned for enterprise on the pricing page.

Features

  • audio-out streaming over WebSocket
  • character and phoneme timestamps
  • voice cloning
  • PerTh watermarking
  • deepfake detection products

Pricing

WhatPriceUnitNotes
TTSNot published on resemble.ai/pricing (detection-only page)Third-party figures conflict: $0.0005/s (Flex) vs $0.006/s (older PAYG)
How the per-minute estimate was worked out

Not verifiable from official pages; third-party sources range from ~$0.03 to ~$0.36 per minute

Free tier: Flex pay-as-you-go credits never expire (detection page); TTS free tier not verified

Source: resemble.ai

Setup

  1. Get a Business-tier (or higher) API key.
  2. Create or pick a voice and project.
  3. Open wss://websocket.cluster.resemble.ai/stream and send one JSON job per utterance.
  4. Read audio frames until audio_end.

Endpoint

wss://websocket.cluster.resemble.ai/stream

Authentication

API key (Bearer) per docs

Warnings

No incremental text input

The WebSocket takes a complete text/SSML payload per request and streams audio back. For LLM output you must chunk sentences yourself.

WebSocket requires Business plan

Lower plans get Unauthorized on the streaming socket.

TTS pricing is not public

resemble.ai/pricing lists only detection products. Get a written quote before committing.

Watermarked output

Chatterbox output carries Resemble's imperceptible PerTh watermark; fine for most uses but note it if you post-process audio.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • WebSocket available on Business plans and above
  • Default 20 simultaneous sessions cluster-wide and 20 parallel connections per key
  • 3,000 characters per request

Models and products

NameStatusNotes
chatterbox-turboGAUsed in the WebSocket docs example; English; vendor says sub-200 ms.
chatterbox (multilingual)GAOpen-weight family also available (MIT).

Docs and sources

Docs

Sources used

Not fully verified

TTS prices; auth header specifics; full model list.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?