GA Inworld AI

Inworld TTS (Realtime TTS-2, TTS-2 Flash)

Low-cost, fast TTS with natural-language steering on TTS-2 and a very fast Flash variant. Bidirectional WebSocket with multiple contexts, timestamps, cloning and an OpenAI-compatible endpoint.

Est. per minute$0.0063 - 0.022

Overview

Best for: Cost-sensitive, high-volume voice agents and games that still want cloning and timestamps.

At a glance

$/1M chars$25
Free tierYes
CloningYes
Text stream inYes
TimestampsYes
EmotionYes
8 kHz phoneYes
Latency ms100
Languages15
Concurrency5
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

Figures are TTS-2 (P90 server TTFB 100 ms, $25/1M On-Demand). TTS-2 Flash: 20 ms P90, $15/1M, ignores steering. Languages: 15 production quality per release notes; docs claim 200+. Concurrency 5 On-Demand (Creator 10, Builder 50). Free start includes up to 70 minutes. Zero data retention supported.

Audio in

Text with inline tags (pauses, pronunciation, steering on TTS-2)

Audio out

LINEAR16 (WAV header per chunk), PCM, MP3 (default), OGG_OPUS, ALAW, MULAW, WAV; 8-48 kHz (default 48 kHz)

Languages

Docs say 200+ languages and locales; release notes describe 15 production-quality plus 90+ experimental for TTS-2

Voices

Catalog voices; instant cloning on all plans (100 custom voices on On-Demand); professional cloning (beta) on TTS-2

Latency

Vendor: P90 server-side TTFB 100 ms (TTS-2), 20 ms (TTS-2 Flash).

Regions

Not specified

Compliance

Zero data retention supported (models page). Certifications not re-verified.

Features

  • input streaming (bidirectional WebSocket)
  • multiple contexts per connection
  • autoMode sentence buffering
  • word timestamps, phonemes and visemes (TTS-2)
  • voice cloning
  • voice design
  • zero data retention
  • OpenAI-compatible POST /v1/audio/speech (2026-09-11)

Pricing

WhatPriceUnitNotes
TTS-2 Flash, On-Demand$15per 1M charactersFree to start, up to 70 min TTS included
TTS-2, On-Demand$25per 1M characters
Creator $25/mo$20 / $10per 1M characters (TTS-2 / Flash)Concurrency 10
Builder $100/mo$17.50 / $9per 1M charactersConcurrency 50
Developer $300/mo$15 / $8per 1M charactersConcurrency 150
Growth $1,500/mo$12.50 / $7per 1M charactersConcurrency 500
EnterpriseAs low as $5 (TTS-2), sub-$5 (Flash)per 1M charactersCustom
How the per-minute estimate was worked out

900 chars/min. Low = TTS-2 Flash on Growth ($7/1M); high = TTS-2 on On-Demand ($25/1M).

Free tier: On-Demand: up to 70 minutes of TTS

Source: inworld.ai

Setup

  1. Get an API key (Basic credential) from the Inworld portal.
  2. Connect to wss://api.inworld.ai/tts/v1/voice:streamBidirectional with Authorization: Basic <key>.
  3. Send create (voice_id, model_id, audio_config), then send_text messages, then flush_context / close_context.
  4. Decode result.audioChunk.audioContent (base64).

Endpoint

wss://api.inworld.ai/tts/v1/voice:streamBidirectional

Authentication

Authorization: Basic <INWORLD_API_KEY> (the key from the portal is already Base64)

Quick start javascript

// npm i ws  - message shapes follow Inworld's official example_websocket.js
import WebSocket from "ws";
import fs from "fs";

const ws = new WebSocket("wss://api.inworld.ai/tts/v1/voice:streamBidirectional", {
  headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` },
});
const out = fs.createWriteStream("out_24k_s16le.pcm");
const ctx = "turn-1";

ws.on("open", () => {
  ws.send(JSON.stringify({ context_id: ctx, create: {
    voice_id: "Ashley", model_id: "inworld-tts-2-flash",
    audio_config: { audio_encoding: "PCM", sample_rate_hertz: 24000 } } }));
  for (const t of ["Hello there. ", "Streaming text ", "from an LLM."])
    ws.send(JSON.stringify({ context_id: ctx, send_text: { text: t } }));
  ws.send(JSON.stringify({ context_id: ctx, close_context: {} }));
});
ws.on("message", (raw) => {
  const r = JSON.parse(raw.toString()).result;
  const b64 = r?.audioChunk?.audioContent || r?.audioContent;
  if (b64) out.write(Buffer.from(b64, "base64"));
  if (r?.contextClosed) ws.close();
});
ws.on("close", () => out.end());

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Steering is ignored on Flash

inworld-tts-2-flash ignores the instruction field and [shouting]-style tags. Use inworld-tts-2 if you rely on prompt-based delivery.

LINEAR16 puts a WAV header in every chunk

With audio_encoding LINEAR16 each chunk carries a WAV header, which clicks if concatenated. Use PCM (headerless) for streaming playback.

Characters counted in UTF-16 code units

Limits and billing count UTF-16 code units, so many emoji and some scripts count as two. Strip emoji from LLM output.

Language claims vary

Docs say 200+ languages while release notes say 15 production-quality plus 90+ experimental. Test non-English quality before committing.

Steering persistence changed

Since 2026-08-06 steering instructions persist until changed or [reset]; older code that assumed per-request tags may behave differently.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Concurrent requests: On-Demand 5, Creator 10, Builder 50, Developer 150, Growth 500
  • 2,000 UTF-16 code units per send_text message
  • Socket closes after 10 min of inactivity across all contexts
  • Sync 2,000 chars, HTTP streaming 4,000 chars, async 100,000 chars per request

Models and products

NameStatusNotes
inworld-tts-2GA (2026-05-05)Steering via natural language and inline tags; P90 server-side TTFB 100 ms (vendor).
inworld-tts-2-flashGA (2026-08-09)P90 server-side TTFB 20 ms (vendor); steering ignored; cheapest.
inworld-tts-1.5-max / inworld-tts-1.5-miniLegacy (Jan 2026)Still usable per integration docs.

Docs and sources

Docs

Sources used

Not fully verified

Voice name 'Ashley' availability on TTS-2 Flash; per-plan WebSocket connection caps.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?