Resemble AI TTS (Chatterbox models)
Voice cloning company whose hosted API serves its Chatterbox models (e.g. chatterbox-turbo). Its WebSocket streams audio out but takes a whole text/SSML payload per request. The public pricing page now covers deepfake detection only.
Overview
Best for: Teams that want hosted Chatterbox with cloning plus deepfake detection/watermarking from one vendor.
At a glance
Chatterbox Turbo (English) sub-200 ms vendor claim; Chatterbox Multilingual open weights cover 23+ languages. TTS pricing not published. WebSocket needs the Business plan; 20 sessions is the default there. Timestamps are character and phoneme level. SOC 2 documentation mentioned for enterprise.
Text or SSML, max 3,000 characters excluding tags
wav/pcm etc.; example uses wav 32 kHz PCM_32; JSON (base64) or binary frames
English for Turbo; Chatterbox Multilingual covers 23+ languages
Voice cloning (rapid and professional)
Vendor (Chatterbox README): hosted service sub-200 ms.
Not specified
Enterprise on-prem available; SOC 2 documentation mentioned for enterprise on the pricing page.
Features
- audio-out streaming over WebSocket
- character and phoneme timestamps
- voice cloning
- PerTh watermarking
- deepfake detection products
Pricing
| What | Price | Unit |
|---|---|---|
| TTS | Not published on resemble.ai/pricing (detection-only page) |
Not verifiable from official pages; third-party sources range from ~$0.03 to ~$0.36 per minute
Free tier: Flex pay-as-you-go credits never expire (detection page); TTS free tier not verified
Source: resemble.ai
Setup
- Get a Business-tier (or higher) API key.
- Create or pick a voice and project.
- Open wss://websocket.cluster.resemble.ai/stream and send one JSON job per utterance.
- Read audio frames until audio_end.
Endpoint
wss://websocket.cluster.resemble.ai/stream
Authentication
API key (Bearer) per docs
Warnings
No incremental text input
The WebSocket takes a complete text/SSML payload per request and streams audio back. For LLM output you must chunk sentences yourself.
WebSocket requires Business plan
Lower plans get Unauthorized on the streaming socket.
TTS pricing is not public
resemble.ai/pricing lists only detection products. Get a written quote before committing.
Watermarked output
Chatterbox output carries Resemble's imperceptible PerTh watermark; fine for most uses but note it if you post-process audio.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- WebSocket available on Business plans and above
- Default 20 simultaneous sessions cluster-wide and 20 parallel connections per key
- 3,000 characters per request
Models and products
| Name | Status |
|---|---|
| chatterbox-turbo | GA |
| chatterbox (multilingual) | GA |
Docs and sources
Docs
Sources used
- resemble.ai/pricing/
- docs.resemble.ai/voice-generation/text-to-speech/streaming-websocket
- voiceflow.com/blog/resemble-ai.md
- github.com/resemble-ai/chatterbox
TTS prices; auth header specifics; full model list.