GA xAI

xAI Grok Voice Agent API

Realtime speech-to-speech WebSocket API for Grok voice models with built-in web search, X search, file search and MCP tools, billed at a flat per-minute rate. Largely OpenAI Realtime compatible, so it is an easy second provider.

Est. per minute$0.08
1 high-severity warning

Overview

Best for: Teams wanting simple flat per-minute pricing, built-in live web and X search, voice cloning, or an OpenAI-compatible fallback provider.

At a glance

Flat $/min$0.08
Free tierNo
Native S2SYes
ToolsYes
Own LLMNo
CloningYes
WebRTCNo
WebSocketYes
Phone / SIPNo
HIPAAYes
SOC 2Yes
EU dataYes
Open weightsNo

Vendor claims sub-second latency (no ms figure). No language list for the voice agent (TTS 20, STT 38+). Compliance items are vendor claims. Tool calls and text inputs billed extra. Telephony via partners (Twilio, Voximplant).

Audio in

audio/pcm at 8000, 16000, 22050, 24000 (default), 32000, 44100 or 48000 Hz; audio/pcmu, audio/pcma (G.711); audio/opus. JSON base64 or binary transport.

Audio out

Same format options as input; speed 0.7 to 1.5.

Languages

Docs say every voice can speak every supported language; TTS lists 20 languages and STT 38+. No explicit list for the voice agent.

Voices

Built-in voices including eve (default), ara and rex (full list via GET /v1/tts/voices); custom voices cloned from a reference clip up to 120 s.

Latency

Vendor claim: sub-second latency.

Regions

us-east-1, eu-west-1, us-saltlake-2; EU data residency options per xAI docs.

Compliance

xAI states SOC 2 Type II, HIPAA eligible with a BAA, GDPR with EU data residency options, and that audio is never stored or used for training (vendor claims).

Features

  • server_vad turn detection with optional idle_timeout_ms re-engagement
  • reasoning.effort high or none
  • tools: web_search, x_search, file_search (Collections), remote MCP, custom functions
  • custom cloned voices
  • session resumption (opt-in conversation caching)
  • replace (spoken substitutions before TTS), keyterms and language hints
  • binary audio transport
  • OpenAI Realtime client compatibility with a base URL change
  • partner integrations: LiveKit, Pipecat, Twilio, Voximplant

Pricing

WhatPriceUnitNotes
Speech to speech (grok-voice-think-fast-2.0)$0.08per minute ($4.80 per hour)
Text input during a voice session$0.004per text inputthird-party sources describe it as a flat fee per conversation.item.create event
web_search tool$5per 1,000 callsxAI tools price list; application to voice sessions not stated
x_search tool$5 per 1,000 posts, $10 per 1,000 profiles
Speech to text (separate API)$0.10 REST / $0.20 streamingper hour
Text to speech (separate API)$15.00per 1M characters
How the per-minute estimate was worked out

Flat $0.08 per minute on grok-voice-think-fast-2.0 regardless of context length. Tool calls and text inputs are extra. xAI does not state whether minutes are session wall-clock time or audio time, so assume wall-clock time including silence.

Audio token rate

Not applicable (billed per minute).

Free tier: None documented for the voice agent.

Source: docs.x.ai

Setup

  1. Create an xAI account at console.x.ai, add billing credit and create an API key.
  2. Server: connect to the realtime WebSocket with Authorization: Bearer <key> and send session.update.
  3. Browser: your server calls POST https://api.x.ai/v1/realtime/client_secrets with {"expires_after": {"seconds": 300}} and returns the token; the browser opens the WebSocket with subprotocol xai-client-secret.<token>.
  4. Stream input_audio_buffer.append and play response.output_audio.delta.
  5. Pin grok-voice-think-fast-2.0 (or 1.0) instead of grok-voice-latest in production.

Endpoint

wss://api.x.ai/v1/realtime?model=grok-voice-latest

Authentication

API key as Authorization: Bearer on the server. Browsers use short-lived tokens from POST https://api.x.ai/v1/realtime/client_secrets passed in the Sec-WebSocket-Protocol header with the xai-client-secret. prefix.

Quick start javascript

import WebSocket from "ws"; // server side; browsers use an ephemeral token via the sec-websocket-protocol header
const ws = new WebSocket("wss://api.x.ai/v1/realtime?model=grok-voice-latest", {
  headers: { Authorization: `Bearer ${process.env.XAI_API_KEY}` },
});
ws.on("open", () => ws.send(JSON.stringify({
  type: "session.update",
  session: {
    voice: "eve",
    instructions: "You are a friendly support agent. Keep answers short.",
    turn_detection: { type: "server_vad" },
    audio: {
      input: { format: { type: "audio/pcm", rate: 24000 } },
      output: { format: { type: "audio/pcm", rate: 24000 } },
    },
    tools: [{ type: "web_search" }],
  },
})));
export function sendAudio(pcm16) { // 24 kHz mono PCM16
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "response.output_audio.delta" || ev.type === "response.audio.delta")
    playPcm16(Buffer.from(ev.delta, "base64"));
  if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
  if (ev.type === "error") console.error(ev.error);
});

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Alias moved and price rose

grok-voice-latest now routes to think-fast-2.0 at $0.08/min, up from $0.05/min on 1.0 (per third-party reports). Pin the model id you tested and priced.

Per-minute billing includes idle time

Flat per-minute billing means open but silent sessions likely cost money. Close sockets on hang-up and use idle timeouts.

Not fully OpenAI compatible

Transcription delta is conversation.item.input_audio_transcription.updated with a cumulative transcript; conversation.item.retrieve, rate_limits.updated and conversation.item.done are unsupported. Test your OpenAI client code paths.

Tool fees add up

Built-in web_search and x_search are convenient but priced per call or per result on the general tools price list. Cap tool usage in instructions and log it.

Thin published limits

No public session length, concurrency or rate-limit numbers. Load test and ask xAI sales for written limits before a launch.

Default reasoning is high

reasoning.effort defaults to high; set none for snappier small talk if latency matters.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Session duration and concurrency limits are not published
  • Resumption history dropped after 30 minutes of inactivity
  • keyterms: max 100 terms, 50 characters each
  • Ephemeral client secrets: expiry set by expires_after.seconds (example 300); session and anchor fields not supported

Models and products

NameStatusNotes
grok-voice-think-fast-2.0GACurrent speech-to-speech model; alias grok-voice-latest points to it. Regions us-east-1, eu-west-1, us-saltlake-2.
grok-voice-think-fast-1.0GAPrevious model ($0.05/min per third-party reports); must be pinned explicitly since the alias moved to 2.0 (reported Aug 5, 2026).

Docs and sources

Docs

Sources used

Not fully verified

How billable minutes are counted; whether tool fees apply inside voice sessions; session length and concurrency limits; 1.0 pricing and alias switch date (third-party); full voice and language lists.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?