xAI Grok Voice Agent API
Realtime speech-to-speech WebSocket API for Grok voice models with built-in web search, X search, file search and MCP tools, billed at a flat per-minute rate. Largely OpenAI Realtime compatible, so it is an easy second provider.
Overview
Best for: Teams wanting simple flat per-minute pricing, built-in live web and X search, voice cloning, or an OpenAI-compatible fallback provider.
At a glance
Vendor claims sub-second latency (no ms figure). No language list for the voice agent (TTS 20, STT 38+). Compliance items are vendor claims. Tool calls and text inputs billed extra. Telephony via partners (Twilio, Voximplant).
audio/pcm at 8000, 16000, 22050, 24000 (default), 32000, 44100 or 48000 Hz; audio/pcmu, audio/pcma (G.711); audio/opus. JSON base64 or binary transport.
Same format options as input; speed 0.7 to 1.5.
Docs say every voice can speak every supported language; TTS lists 20 languages and STT 38+. No explicit list for the voice agent.
Built-in voices including eve (default), ara and rex (full list via GET /v1/tts/voices); custom voices cloned from a reference clip up to 120 s.
Vendor claim: sub-second latency.
us-east-1, eu-west-1, us-saltlake-2; EU data residency options per xAI docs.
xAI states SOC 2 Type II, HIPAA eligible with a BAA, GDPR with EU data residency options, and that audio is never stored or used for training (vendor claims).
Features
- server_vad turn detection with optional idle_timeout_ms re-engagement
- reasoning.effort high or none
- tools: web_search, x_search, file_search (Collections), remote MCP, custom functions
- custom cloned voices
- session resumption (opt-in conversation caching)
- replace (spoken substitutions before TTS), keyterms and language hints
- binary audio transport
- OpenAI Realtime client compatibility with a base URL change
- partner integrations: LiveKit, Pipecat, Twilio, Voximplant
Pricing
| What | Price | Unit |
|---|---|---|
| Speech to speech (grok-voice-think-fast-2.0) | $0.08 | per minute ($4.80 per hour) |
| Text input during a voice session | $0.004 | per text input |
| web_search tool | $5 | per 1,000 calls |
| x_search tool | $5 per 1,000 posts, $10 per 1,000 profiles | |
| Speech to text (separate API) | $0.10 REST / $0.20 streaming | per hour |
| Text to speech (separate API) | $15.00 | per 1M characters |
Flat $0.08 per minute on grok-voice-think-fast-2.0 regardless of context length. Tool calls and text inputs are extra. xAI does not state whether minutes are session wall-clock time or audio time, so assume wall-clock time including silence.
Not applicable (billed per minute).
Free tier: None documented for the voice agent.
Source: docs.x.ai
Setup
- Create an xAI account at console.x.ai, add billing credit and create an API key.
- Server: connect to the realtime WebSocket with Authorization: Bearer <key> and send session.update.
- Browser: your server calls POST https://api.x.ai/v1/realtime/client_secrets with {"expires_after": {"seconds": 300}} and returns the token; the browser opens the WebSocket with subprotocol xai-client-secret.<token>.
- Stream input_audio_buffer.append and play response.output_audio.delta.
- Pin grok-voice-think-fast-2.0 (or 1.0) instead of grok-voice-latest in production.
Endpoint
wss://api.x.ai/v1/realtime?model=grok-voice-latest
Authentication
API key as Authorization: Bearer on the server. Browsers use short-lived tokens from POST https://api.x.ai/v1/realtime/client_secrets passed in the Sec-WebSocket-Protocol header with the xai-client-secret. prefix.
Quick start javascript
import WebSocket from "ws"; // server side; browsers use an ephemeral token via the sec-websocket-protocol header
const ws = new WebSocket("wss://api.x.ai/v1/realtime?model=grok-voice-latest", {
headers: { Authorization: `Bearer ${process.env.XAI_API_KEY}` },
});
ws.on("open", () => ws.send(JSON.stringify({
type: "session.update",
session: {
voice: "eve",
instructions: "You are a friendly support agent. Keep answers short.",
turn_detection: { type: "server_vad" },
audio: {
input: { format: { type: "audio/pcm", rate: 24000 } },
output: { format: { type: "audio/pcm", rate: 24000 } },
},
tools: [{ type: "web_search" }],
},
})));
export function sendAudio(pcm16) { // 24 kHz mono PCM16
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.output_audio.delta" || ev.type === "response.audio.delta")
playPcm16(Buffer.from(ev.delta, "base64"));
if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
if (ev.type === "error") console.error(ev.error);
});
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Alias moved and price rose
grok-voice-latest now routes to think-fast-2.0 at $0.08/min, up from $0.05/min on 1.0 (per third-party reports). Pin the model id you tested and priced.
Per-minute billing includes idle time
Flat per-minute billing means open but silent sessions likely cost money. Close sockets on hang-up and use idle timeouts.
Not fully OpenAI compatible
Transcription delta is conversation.item.input_audio_transcription.updated with a cumulative transcript; conversation.item.retrieve, rate_limits.updated and conversation.item.done are unsupported. Test your OpenAI client code paths.
Tool fees add up
Built-in web_search and x_search are convenient but priced per call or per result on the general tools price list. Cap tool usage in instructions and log it.
Thin published limits
No public session length, concurrency or rate-limit numbers. Load test and ask xAI sales for written limits before a launch.
Default reasoning is high
reasoning.effort defaults to high; set none for snappier small talk if latency matters.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Session duration and concurrency limits are not published
- Resumption history dropped after 30 minutes of inactivity
- keyterms: max 100 terms, 50 characters each
- Ephemeral client secrets: expiry set by expires_after.seconds (example 300); session and anchor fields not supported
Models and products
| Name | Status |
|---|---|
| grok-voice-think-fast-2.0 | GA |
| grok-voice-think-fast-1.0 | GA |
Docs and sources
Docs
Sources used
- docs.x.ai/docs/guides/voice
- docs.x.ai/developers/model-capabilities/audio/voice-agent
- docs.x.ai/developers/model-capabilities/audio/ephemeral-tokens
- docs.x.ai/developers/models
- docs.x.ai/developers/pricing
- docs.x.ai/developers/models/grok-voice-think-fast-2.0
- eesel.ai/blog/grok-voice-think-fast-2-pricing
How billable minutes are counted; whether tool fees apply inside voice sessions; session length and concurrency limits; 1.0 pricing and alias switch date (third-party); full voice and language lists.