Inworld Realtime API
An OpenAI-Realtime-compatible voice API that chains Inworld STT, any of 100+ routed LLMs and Inworld TTS-2, over WebSocket or WebRTC. Good if you already speak the OpenAI Realtime protocol but want cheaper voices and free choice of LLM.
Overview
Best for: Games, characters and consumer apps that want OpenAI-Realtime-style integration with cheaper expressive TTS and free LLM choice.
At a glance
Billed per TTS character and STT hour, not per minute; vendor-based estimate about $0.01 to $0.03 per minute before LLM. LLM chosen from a router of 100+ provider models, billed at cost; custom endpoint not stated. Concurrency 5 on On-Demand, 50 on Builder.
PCM16 mono 24 kHz default; G.711 mu-law/A-law 8 kHz; float32
Same options
Depends on STT/TTS models chosen; not listed on the pages read.
Inworld voice library (e.g. Clive) and cloned voices.
Not stated on the pages read.
Not stated.
Not stated on the pages read.
Features
- OpenAI Realtime-style events
- semantic VAD with interrupt_response
- function calling (tools, tool_choice)
- mid-session model/voice switching
- memory, back-channel and responsiveness extensions (providerData)
Pricing
| What | Price | Unit |
|---|---|---|
| Realtime TTS-2 | $25 (On-Demand) down to $12.50 (Growth) | per 1M characters |
| TTS-2 Flash | $15 down to $7 | per 1M characters |
| STT 1 | $0.15 (On-Demand) / $0.10 (paid plans) | per hour |
| LLM | provider cost | per token |
| Plans | On-Demand free, Creator $25, Builder $100, Developer $300, Growth $1,500 | per month |
Own estimate: STT about $0.0025/min plus TTS-2 about $0.0125-0.025 per minute of agent speech (1,000 chars/min), before LLM cost.
Inworld estimates 1 minute of speech is about 1,000 characters of TTS.
Free tier: On-Demand plan free to start with up to 70 TTS minutes and up to 400 STT minutes.
Source: inworld.ai
Setup
- Create an Inworld account and API key in the Portal.
- Server: open the realtime WebSocket with Basic auth; browser: mint a JWT session token and use Bearer auth (or WebRTC).
- On session.created send session.update choosing LLM, STT, TTS model and voice; stream input_audio_buffer.append and play response.output_audio.delta.
Endpoint
wss://api.inworld.ai/api/v1/realtime/session?key=<session-id>&protocol=realtime
Authentication
Server: Authorization: Basic <api-key>; browser: Authorization: Bearer <jwt>
Quick start javascript
// npm i ws (server-side; browsers use a short-lived JWT with Bearer auth)
import WebSocket from "ws";
const sessionId = crypto.randomUUID();
const ws = new WebSocket(
`wss://api.inworld.ai/api/v1/realtime/session?key=${sessionId}&protocol=realtime`,
{ headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` } }
);
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "session.created") {
ws.send(JSON.stringify({
type: "session.update",
session: {
type: "realtime",
model: "openai/gpt-4o-mini", // LLM billed at provider cost
instructions: "You are a friendly narrator.",
output_modalities: ["audio", "text"],
audio: {
input: { transcription: { model: "inworld/inworld-stt-1" },
turn_detection: { type: "semantic_vad", create_response: true, interrupt_response: true } },
output: { voice: "Clive", model: "inworld-tts-2" },
},
},
}));
}
if (ev.type === "response.output_audio.delta") play(Buffer.from(ev.delta, "base64"));
});
// 24 kHz mono PCM16, 60-100 ms chunks (OpenAI-style event)
export function sendPcm(buf) {
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: buf.toString("base64") }));
}
function play(pcm) {}
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
LLM cost is extra
Published prices cover STT and TTS only; the routed LLM is billed at provider cost on top.
OpenAI-compatible, not identical
Event names follow OpenAI Realtime, but Inworld options live in providerData and full drop-in compatibility is not claimed. Test your existing client.
Low concurrency on free tier
On-Demand allows 5 concurrent requests; production needs at least Builder (50).
Cascaded pipeline
STT, LLM and TTS are separate; vocal emotion is not passed to the LLM.
Character-based TTS billing
Verbose LLM output directly increases TTS cost; cap max_output_tokens.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Concurrent requests: 5 (On-Demand), 10 (Creator), 50 (Builder), 150 (Developer), 500 (Growth), custom (Enterprise)
Models and products
| Name | Status |
|---|---|
| STT: inworld/inworld-stt-1 | GA |
| TTS: inworld-tts-2 (and TTS-2 Flash) | GA |
| LLM: provider/model ids or routers (e.g. openai/gpt-4o-mini, inworld/latency-optimizer-ab-test) | GA |
Docs and sources
Docs
Sources used
- docs.inworld.ai/realtime/connect/websocket.md
- docs.inworld.ai/realtime/usage/using-realtime-models.md
- inworld.ai/pricing
Latency, languages, regions, compliance; whether realtime sessions carry any per-minute platform fee beyond STT/TTS/LLM.