GLM-4-Voice-9B
Zhipu's 2024 open end-to-end Chinese/English voice chat model (9B LLM plus speech tokenizer and decoder). Historically important and still usable, but older than the other options here.
Overview
Best for: Chinese-first research baselines.
At a glance
Weights under the custom GLM-4 model licence (code Apache-2.0); check commercial terms. Streaming decoding can start after about 10 speech tokens.
Speech (12.5 tokens per second tokenizer)
Speech via flow-matching decoder
Chinese and English.
Instruction-controlled style (emotion, speed, dialect) per repo; no voice list.
Vendor: generation can start after as few as 10 speech tokens; synthesis needs as few as 20 output tokens.
Self-hosted.
Self-hosted.
Not stated; bf16 9B model plus decoder suggests 24 GB-class GPU, int4 less (own estimate).
Code Apache-2.0; weights under the GLM-4 model licence (custom; read before commercial use).
Features
- end-to-end speech chat
- streaming decoding
- int4 option
Pricing
No public price list.
No licence fee; you pay for GPU time.
n/a (self-hosted)
Free tier: Open weights
Source: github.com
Setup
- CUDA GPU (bf16, or int4 for less VRAM).
- Start the model server, then the web demo with tokenizer and decoder paths.
Endpoint
Local: http://127.0.0.1:8888 (web demo), model server on port 10000
Authentication
n/a
Quick start python
git clone https://github.com/zai-org/GLM-4-Voice && cd GLM-4-Voice
pip install -r requirements.txt
git clone https://huggingface.co/THUDM/glm-4-voice-decoder
python model_server.py --host localhost --model-path THUDM/glm-4-voice-9b --port 10000 --dtype bfloat16 --device cuda:0
python web_demo.py --tokenizer-path THUDM/glm-4-voice-tokenizer --model-path THUDM/glm-4-voice-9b --flow-path ./glm-4-voice-decoder
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Custom weight licence
Weights follow the GLM-4 model agreement, not Apache; commercial terms must be checked.
Dated model
From October 2024; newer open models outperform it.
Three components to wire
LLM, tokenizer and decoder are separate downloads and processes.
EN/ZH only
Only Chinese and English speech.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Released October 2024; not full duplex
Models and products
| Name | Status |
|---|---|
| zai-org/glm-4-voice-9b (also THUDM/glm-4-voice-9b) | GA |
| kyutai/glm-4-voice-of-reason-9b | Preview |
Docs and sources
Docs
Sources used
VRAM; GLM-4 licence commercial terms; requirements install step.