GA Zhipu AI (bigmodel.cn)

GLM-Realtime

Zhipu's realtime voice (and passive video) model over an OpenAI-Realtime-style WebSocket, billed simply per minute. Good for Chinese-market voice and video-call assistants; short memory makes it a poor fit for long sessions.

Est. per minute$0.025 - 0.30
1 high-severity warning

Overview

Best for: Chinese-market voice or video-call assistants with short interactions and predictable per-minute cost.

At a glance

Flat $/min$0.025
Native S2SYes
ToolsYes
Image inYes
Own LLMNo
Voices7
Context tokens8,000
Concurrency5
WebRTCNo
WebSocketYes
Phone / SIPNo
EU dataNo
Open weightsNo

Price is glm-realtime-flash audio, 0.18 CNY/min converted at 7.1 CNY per USD (air 0.3 CNY/min; video 1.2 to 2.1 CNY/min). Concurrency 5 at account tier V0, up to 20 at V3. Context 8K for audio (about 2 min memory), 32K for video. China platform only.

Audio in

wav or pcm (pcm16 = 16 kHz, pcm24 = 24 kHz), mono 16-bit

Audio out

PCM 24 kHz mono 16-bit

Languages

Multilingual with automatic language detection (no list published); replies in the user's language.

Voices

tongtong (default), xiaochen, female-tianmei, female-shaonv, male-qn-daxuesheng, male-qn-jingying, lovely_girl

Latency

Not published.

Regions

Chinese mainland (open.bigmodel.cn). Not confirmed on the international z.ai platform (its GLM-Realtime doc URL returned 404).

Compliance

Not stated on the page read.

Features

  • server VAD or client VAD
  • interruption (interrupt_response, response.cancel)
  • function calling (voice calls only)
  • built-in web search (auto_search)
  • passive video mode (video_passive)
  • near/far-field noise reduction
  • greeting config
  • singing

Pricing

WhatPriceUnitNotes
glm-realtime-flash audio call0.18 CNYper minute
glm-realtime-flash video call1.2 CNYper minute
glm-realtime-air audio call0.3 CNYper minute
glm-realtime-air video call2.1 CNYper minute
How the per-minute estimate was worked out

Conversion at about 7.1 CNY per USD: flash audio about $0.025, air audio about $0.042, air video about $0.30 per minute.

Audio token rate

Not token-billed.

Free tier: Not stated on the GLM-Realtime page.

Source: docs.bigmodel.cn

Setup

  1. Register on open.bigmodel.cn and create an API key (real-name verification may be required).
  2. Connect to the realtime WebSocket with Authorization: Bearer <key> (a JWT also works).
  3. Send session.update with session.model set to glm-realtime-flash or glm-realtime-air.
  4. Stream input_audio_buffer.append (pcm16) and play response.audio.delta (24 kHz PCM).

Endpoint

wss://open.bigmodel.cn/api/paas/v4/realtime

Authentication

Authorization header with API key or JWT

Quick start javascript

// npm i ws
import WebSocket from "ws";

const ws = new WebSocket("wss://open.bigmodel.cn/api/paas/v4/realtime", {
  headers: { Authorization: `Bearer ${process.env.ZHIPU_API_KEY}` },
});

ws.on("open", () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      model: "glm-realtime-flash",        // or glm-realtime-air
      modalities: ["text", "audio"],
      instructions: "You are a helpful voice assistant.",
      voice: "tongtong",
      input_audio_format: "pcm16",        // 16 kHz mono 16-bit
      output_audio_format: "pcm",         // 24 kHz mono 16-bit
      beta_fields: { chat_mode: "audio" },
    },
  }));
});

// ~100 ms frames, at most 50 messages per second
export function sendPcm(buf) {
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: buf.toString("base64") }));
}

ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "response.audio.delta") play(Buffer.from(ev.delta, "base64"));
  if (ev.type === "error") console.error(ev);
});
function play(pcm) {}

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Very short memory

Audio sessions get an 8K context (about 20 turns) and roughly 2 minutes of conversation memory; long calls will forget earlier details. Re-inject key facts via instructions.

China platform only

Documented on open.bigmodel.cn with CNY pricing; international z.ai availability is not confirmed. Expect Chinese account requirements and mainland data processing.

Low concurrency on new accounts

Starting tier allows 5 concurrent sessions; higher tiers depend on account level.

Video is 6-7x the audio price

Video mode is 1.2 to 2.1 CNY per minute and needs at least one image uploaded before creating a response, or it errors.

Short replies

max_response_output_tokens is capped at 1024, so long answers get cut.

Tools only in voice mode

Function calling is documented as voice-call only, not video calls.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Context 8K for audio calls (about 20 turns per docs) and 32K for video calls
  • Conversation memory up to about 2 minutes
  • max_response_output_tokens up to 1024
  • Client VAD mode: max 30 s per upload; send at most 50 messages per second (100 ms frames recommended)
  • Concurrency by account tier: V0 5, V1 10, V2 15, V3 20

Models and products

NameStatusNotes
glm-realtime-flashGA9B model per docs.
glm-realtime-airGA32B model per docs.
glm-realtimeGADefault value of session.model; no separate price listed.

Docs and sources

Docs

Sources used

Not fully verified

International (z.ai) availability, free tier, real-name rules for foreigners, latency.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?