GA Zhipu AI (zai-org)

GLM-4-Voice-9B

Zhipu's 2024 open end-to-end Chinese/English voice chat model (9B LLM plus speech tokenizer and decoder). Historically important and still usable, but older than the other options here.

Est. per minuten/a

Overview

Best for: Chinese-first research baselines.

At a glance

TypeVoice-to-voice
Params B9
LicenceCustom
Full duplexNo
StreamingYes
Languages2

Weights under the custom GLM-4 model licence (code Apache-2.0); check commercial terms. Streaming decoding can start after about 10 speech tokens.

Audio in

Speech (12.5 tokens per second tokenizer)

Audio out

Speech via flow-matching decoder

Languages

Chinese and English.

Voices

Instruction-controlled style (emotion, speed, dialect) per repo; no voice list.

Latency

Vendor: generation can start after as few as 10 speech tokens; synthesis needs as few as 20 output tokens.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

Not stated; bf16 9B model plus decoder suggests 24 GB-class GPU, int4 less (own estimate).

Licence

Code Apache-2.0; weights under the GLM-4 model licence (custom; read before commercial use).

Features

  • end-to-end speech chat
  • streaming decoding
  • int4 option

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: github.com

Setup

  1. CUDA GPU (bf16, or int4 for less VRAM).
  2. Start the model server, then the web demo with tokenizer and decoder paths.

Endpoint

Local: http://127.0.0.1:8888 (web demo), model server on port 10000

Authentication

n/a

Quick start python

git clone https://github.com/zai-org/GLM-4-Voice && cd GLM-4-Voice
pip install -r requirements.txt
git clone https://huggingface.co/THUDM/glm-4-voice-decoder
python model_server.py --host localhost --model-path THUDM/glm-4-voice-9b --port 10000 --dtype bfloat16 --device cuda:0
python web_demo.py --tokenizer-path THUDM/glm-4-voice-tokenizer --model-path THUDM/glm-4-voice-9b --flow-path ./glm-4-voice-decoder

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Custom weight licence

Weights follow the GLM-4 model agreement, not Apache; commercial terms must be checked.

Dated model

From October 2024; newer open models outperform it.

Three components to wire

LLM, tokenizer and decoder are separate downloads and processes.

EN/ZH only

Only Chinese and English speech.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Released October 2024; not full duplex

Models and products

NameStatusNotes
zai-org/glm-4-voice-9b (also THUDM/glm-4-voice-9b)GALLM part; needs glm-4-voice-tokenizer and glm-4-voice-decoder.
kyutai/glm-4-voice-of-reason-9bPreview2026 Kyutai research derivative; not verified.

Docs and sources

Docs

Sources used

Not fully verified

VRAM; GLM-4 licence commercial terms; requirements install step.

Similar open models APIs

Spotted a wrong price or a dead link?