Inworld TTS (Realtime TTS-2, TTS-2 Flash)
Low-cost, fast TTS with natural-language steering on TTS-2 and a very fast Flash variant. Bidirectional WebSocket with multiple contexts, timestamps, cloning and an OpenAI-compatible endpoint.
Overview
Best for: Cost-sensitive, high-volume voice agents and games that still want cloning and timestamps.
At a glance
Figures are TTS-2 (P90 server TTFB 100 ms, $25/1M On-Demand). TTS-2 Flash: 20 ms P90, $15/1M, ignores steering. Languages: 15 production quality per release notes; docs claim 200+. Concurrency 5 On-Demand (Creator 10, Builder 50). Free start includes up to 70 minutes. Zero data retention supported.
Text with inline tags (pauses, pronunciation, steering on TTS-2)
LINEAR16 (WAV header per chunk), PCM, MP3 (default), OGG_OPUS, ALAW, MULAW, WAV; 8-48 kHz (default 48 kHz)
Docs say 200+ languages and locales; release notes describe 15 production-quality plus 90+ experimental for TTS-2
Catalog voices; instant cloning on all plans (100 custom voices on On-Demand); professional cloning (beta) on TTS-2
Vendor: P90 server-side TTFB 100 ms (TTS-2), 20 ms (TTS-2 Flash).
Not specified
Zero data retention supported (models page). Certifications not re-verified.
Features
- input streaming (bidirectional WebSocket)
- multiple contexts per connection
- autoMode sentence buffering
- word timestamps, phonemes and visemes (TTS-2)
- voice cloning
- voice design
- zero data retention
- OpenAI-compatible POST /v1/audio/speech (2026-09-11)
Pricing
| What | Price | Unit |
|---|---|---|
| TTS-2 Flash, On-Demand | $15 | per 1M characters |
| TTS-2, On-Demand | $25 | per 1M characters |
| Creator $25/mo | $20 / $10 | per 1M characters (TTS-2 / Flash) |
| Builder $100/mo | $17.50 / $9 | per 1M characters |
| Developer $300/mo | $15 / $8 | per 1M characters |
| Growth $1,500/mo | $12.50 / $7 | per 1M characters |
| Enterprise | As low as $5 (TTS-2), sub-$5 (Flash) | per 1M characters |
900 chars/min. Low = TTS-2 Flash on Growth ($7/1M); high = TTS-2 on On-Demand ($25/1M).
Free tier: On-Demand: up to 70 minutes of TTS
Source: inworld.ai
Setup
- Get an API key (Basic credential) from the Inworld portal.
- Connect to wss://api.inworld.ai/tts/v1/voice:streamBidirectional with Authorization: Basic <key>.
- Send create (voice_id, model_id, audio_config), then send_text messages, then flush_context / close_context.
- Decode result.audioChunk.audioContent (base64).
Endpoint
wss://api.inworld.ai/tts/v1/voice:streamBidirectional
Authentication
Authorization: Basic <INWORLD_API_KEY> (the key from the portal is already Base64)
Quick start javascript
// npm i ws - message shapes follow Inworld's official example_websocket.js
import WebSocket from "ws";
import fs from "fs";
const ws = new WebSocket("wss://api.inworld.ai/tts/v1/voice:streamBidirectional", {
headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` },
});
const out = fs.createWriteStream("out_24k_s16le.pcm");
const ctx = "turn-1";
ws.on("open", () => {
ws.send(JSON.stringify({ context_id: ctx, create: {
voice_id: "Ashley", model_id: "inworld-tts-2-flash",
audio_config: { audio_encoding: "PCM", sample_rate_hertz: 24000 } } }));
for (const t of ["Hello there. ", "Streaming text ", "from an LLM."])
ws.send(JSON.stringify({ context_id: ctx, send_text: { text: t } }));
ws.send(JSON.stringify({ context_id: ctx, close_context: {} }));
});
ws.on("message", (raw) => {
const r = JSON.parse(raw.toString()).result;
const b64 = r?.audioChunk?.audioContent || r?.audioContent;
if (b64) out.write(Buffer.from(b64, "base64"));
if (r?.contextClosed) ws.close();
});
ws.on("close", () => out.end());
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Steering is ignored on Flash
inworld-tts-2-flash ignores the instruction field and [shouting]-style tags. Use inworld-tts-2 if you rely on prompt-based delivery.
LINEAR16 puts a WAV header in every chunk
With audio_encoding LINEAR16 each chunk carries a WAV header, which clicks if concatenated. Use PCM (headerless) for streaming playback.
Characters counted in UTF-16 code units
Limits and billing count UTF-16 code units, so many emoji and some scripts count as two. Strip emoji from LLM output.
Language claims vary
Docs say 200+ languages while release notes say 15 production-quality plus 90+ experimental. Test non-English quality before committing.
Steering persistence changed
Since 2026-08-06 steering instructions persist until changed or [reset]; older code that assumed per-request tags may behave differently.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Concurrent requests: On-Demand 5, Creator 10, Builder 50, Developer 150, Growth 500
- 2,000 UTF-16 code units per send_text message
- Socket closes after 10 min of inactivity across all contexts
- Sync 2,000 chars, HTTP streaming 4,000 chars, async 100,000 chars per request
Models and products
| Name | Status |
|---|---|
| inworld-tts-2 | GA (2026-05-05) |
| inworld-tts-2-flash | GA (2026-08-09) |
| inworld-tts-1.5-max / inworld-tts-1.5-mini | Legacy (Jan 2026) |
Docs and sources
Docs
Sources used
- inworld.ai/pricing
- docs.inworld.ai/tts/tts-models.md
- docs.inworld.ai/release-notes/tts
- docs.inworld.ai/api-reference/ttsAPI/texttospeech/synthesize-speech-websocket.md
- docs.inworld.ai/tts/synthesize-speech.md
- github.com/inworld-ai/inworld-api-examples/blob/main/tts/js/example_websocket.js
Voice name 'Ashley' availability on TTS-2 Flash; per-plan WebSocket connection caps.