GLM-Realtime
Zhipu's realtime voice (and passive video) model over an OpenAI-Realtime-style WebSocket, billed simply per minute. Good for Chinese-market voice and video-call assistants; short memory makes it a poor fit for long sessions.
Overview
Best for: Chinese-market voice or video-call assistants with short interactions and predictable per-minute cost.
At a glance
Price is glm-realtime-flash audio, 0.18 CNY/min converted at 7.1 CNY per USD (air 0.3 CNY/min; video 1.2 to 2.1 CNY/min). Concurrency 5 at account tier V0, up to 20 at V3. Context 8K for audio (about 2 min memory), 32K for video. China platform only.
wav or pcm (pcm16 = 16 kHz, pcm24 = 24 kHz), mono 16-bit
PCM 24 kHz mono 16-bit
Multilingual with automatic language detection (no list published); replies in the user's language.
tongtong (default), xiaochen, female-tianmei, female-shaonv, male-qn-daxuesheng, male-qn-jingying, lovely_girl
Not published.
Chinese mainland (open.bigmodel.cn). Not confirmed on the international z.ai platform (its GLM-Realtime doc URL returned 404).
Not stated on the page read.
Features
- server VAD or client VAD
- interruption (interrupt_response, response.cancel)
- function calling (voice calls only)
- built-in web search (auto_search)
- passive video mode (video_passive)
- near/far-field noise reduction
- greeting config
- singing
Pricing
| What | Price | Unit |
|---|---|---|
| glm-realtime-flash audio call | 0.18 CNY | per minute |
| glm-realtime-flash video call | 1.2 CNY | per minute |
| glm-realtime-air audio call | 0.3 CNY | per minute |
| glm-realtime-air video call | 2.1 CNY | per minute |
Conversion at about 7.1 CNY per USD: flash audio about $0.025, air audio about $0.042, air video about $0.30 per minute.
Not token-billed.
Free tier: Not stated on the GLM-Realtime page.
Source: docs.bigmodel.cn
Setup
- Register on open.bigmodel.cn and create an API key (real-name verification may be required).
- Connect to the realtime WebSocket with Authorization: Bearer <key> (a JWT also works).
- Send session.update with session.model set to glm-realtime-flash or glm-realtime-air.
- Stream input_audio_buffer.append (pcm16) and play response.audio.delta (24 kHz PCM).
Endpoint
wss://open.bigmodel.cn/api/paas/v4/realtime
Authentication
Authorization header with API key or JWT
Quick start javascript
// npm i ws
import WebSocket from "ws";
const ws = new WebSocket("wss://open.bigmodel.cn/api/paas/v4/realtime", {
headers: { Authorization: `Bearer ${process.env.ZHIPU_API_KEY}` },
});
ws.on("open", () => {
ws.send(JSON.stringify({
type: "session.update",
session: {
model: "glm-realtime-flash", // or glm-realtime-air
modalities: ["text", "audio"],
instructions: "You are a helpful voice assistant.",
voice: "tongtong",
input_audio_format: "pcm16", // 16 kHz mono 16-bit
output_audio_format: "pcm", // 24 kHz mono 16-bit
beta_fields: { chat_mode: "audio" },
},
}));
});
// ~100 ms frames, at most 50 messages per second
export function sendPcm(buf) {
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: buf.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.audio.delta") play(Buffer.from(ev.delta, "base64"));
if (ev.type === "error") console.error(ev);
});
function play(pcm) {}
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Very short memory
Audio sessions get an 8K context (about 20 turns) and roughly 2 minutes of conversation memory; long calls will forget earlier details. Re-inject key facts via instructions.
China platform only
Documented on open.bigmodel.cn with CNY pricing; international z.ai availability is not confirmed. Expect Chinese account requirements and mainland data processing.
Low concurrency on new accounts
Starting tier allows 5 concurrent sessions; higher tiers depend on account level.
Video is 6-7x the audio price
Video mode is 1.2 to 2.1 CNY per minute and needs at least one image uploaded before creating a response, or it errors.
Short replies
max_response_output_tokens is capped at 1024, so long answers get cut.
Tools only in voice mode
Function calling is documented as voice-call only, not video calls.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Context 8K for audio calls (about 20 turns per docs) and 32K for video calls
- Conversation memory up to about 2 minutes
- max_response_output_tokens up to 1024
- Client VAD mode: max 30 s per upload; send at most 50 messages per second (100 ms frames recommended)
- Concurrency by account tier: V0 5, V1 10, V2 15, V3 20
Models and products
| Name | Status |
|---|---|
| glm-realtime-flash | GA |
| glm-realtime-air | GA |
| glm-realtime | GA |
Docs and sources
Docs
Sources used
International (z.ai) availability, free tier, real-name rules for foreigners, latency.