GA Inworld AI

Inworld Realtime API

An OpenAI-Realtime-compatible voice API that chains Inworld STT, any of 100+ routed LLMs and Inworld TTS-2, over WebSocket or WebRTC. Good if you already speak the OpenAI Realtime protocol but want cheaper voices and free choice of LLM.

Est. per minute$0.01 - 0.03
1 high-severity warning

Overview

Best for: Games, characters and consumer apps that want OpenAI-Realtime-style integration with cheaper expressive TTS and free LLM choice.

At a glance

Free tierYes
Native S2SNo
ToolsYes
CloningYes
Concurrency5
WebRTCYes
WebSocketYes
Phone / SIPNo
Open weightsNo

Billed per TTS character and STT hour, not per minute; vendor-based estimate about $0.01 to $0.03 per minute before LLM. LLM chosen from a router of 100+ provider models, billed at cost; custom endpoint not stated. Concurrency 5 on On-Demand, 50 on Builder.

Audio in

PCM16 mono 24 kHz default; G.711 mu-law/A-law 8 kHz; float32

Audio out

Same options

Languages

Depends on STT/TTS models chosen; not listed on the pages read.

Voices

Inworld voice library (e.g. Clive) and cloned voices.

Latency

Not stated on the pages read.

Regions

Not stated.

Compliance

Not stated on the pages read.

Features

  • OpenAI Realtime-style events
  • semantic VAD with interrupt_response
  • function calling (tools, tool_choice)
  • mid-session model/voice switching
  • memory, back-channel and responsiveness extensions (providerData)

Pricing

WhatPriceUnitNotes
Realtime TTS-2$25 (On-Demand) down to $12.50 (Growth)per 1M charactersEnterprise as low as $5.
TTS-2 Flash$15 down to $7per 1M characters
STT 1$0.15 (On-Demand) / $0.10 (paid plans)per hour
LLMprovider costper tokenBilled at cost, separately.
PlansOn-Demand free, Creator $25, Builder $100, Developer $300, Growth $1,500per monthPaid plans include credits equal to the plan price.
How the per-minute estimate was worked out

Own estimate: STT about $0.0025/min plus TTS-2 about $0.0125-0.025 per minute of agent speech (1,000 chars/min), before LLM cost.

Audio token rate

Inworld estimates 1 minute of speech is about 1,000 characters of TTS.

Free tier: On-Demand plan free to start with up to 70 TTS minutes and up to 400 STT minutes.

Source: inworld.ai

Setup

  1. Create an Inworld account and API key in the Portal.
  2. Server: open the realtime WebSocket with Basic auth; browser: mint a JWT session token and use Bearer auth (or WebRTC).
  3. On session.created send session.update choosing LLM, STT, TTS model and voice; stream input_audio_buffer.append and play response.output_audio.delta.

Endpoint

wss://api.inworld.ai/api/v1/realtime/session?key=<session-id>&protocol=realtime

Authentication

Server: Authorization: Basic <api-key>; browser: Authorization: Bearer <jwt>

Quick start javascript

// npm i ws   (server-side; browsers use a short-lived JWT with Bearer auth)
import WebSocket from "ws";

const sessionId = crypto.randomUUID();
const ws = new WebSocket(
  `wss://api.inworld.ai/api/v1/realtime/session?key=${sessionId}&protocol=realtime`,
  { headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` } }
);

ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "session.created") {
    ws.send(JSON.stringify({
      type: "session.update",
      session: {
        type: "realtime",
        model: "openai/gpt-4o-mini",          // LLM billed at provider cost
        instructions: "You are a friendly narrator.",
        output_modalities: ["audio", "text"],
        audio: {
          input: { transcription: { model: "inworld/inworld-stt-1" },
                   turn_detection: { type: "semantic_vad", create_response: true, interrupt_response: true } },
          output: { voice: "Clive", model: "inworld-tts-2" },
        },
      },
    }));
  }
  if (ev.type === "response.output_audio.delta") play(Buffer.from(ev.delta, "base64"));
});

// 24 kHz mono PCM16, 60-100 ms chunks (OpenAI-style event)
export function sendPcm(buf) {
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: buf.toString("base64") }));
}
function play(pcm) {}

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

LLM cost is extra

Published prices cover STT and TTS only; the routed LLM is billed at provider cost on top.

OpenAI-compatible, not identical

Event names follow OpenAI Realtime, but Inworld options live in providerData and full drop-in compatibility is not claimed. Test your existing client.

Low concurrency on free tier

On-Demand allows 5 concurrent requests; production needs at least Builder (50).

Cascaded pipeline

STT, LLM and TTS are separate; vocal emotion is not passed to the LLM.

Character-based TTS billing

Verbose LLM output directly increases TTS cost; cap max_output_tokens.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Concurrent requests: 5 (On-Demand), 10 (Creator), 50 (Builder), 150 (Developer), 500 (Growth), custom (Enterprise)

Models and products

NameStatusNotes
STT: inworld/inworld-stt-1GADefault realtime STT with turn-taking controls.
TTS: inworld-tts-2 (and TTS-2 Flash)GA
LLM: provider/model ids or routers (e.g. openai/gpt-4o-mini, inworld/latency-optimizer-ab-test)GARouter offers 100+ models billed at provider cost.

Docs and sources

Docs

Sources used

Not fully verified

Latency, languages, regions, compliance; whether realtime sessions carry any per-minute platform fee beyond STT/TTS/LLM.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?