MiniCPM-o 4.5
A 9B open omni model that does full-duplex speech (and video) streaming and runs on a single 12-24 GB GPU or a Mac via llama.cpp. The most practical open full-duplex model for English/Chinese with vision.
Overview
Best for: On-prem or edge full-duplex voice plus vision assistants on a single mid-range GPU or Mac.
At a glance
600 ms is vendor time to first token. 11 GB is int4 (bf16 about 19 GB; llama.cpp full duplex needs 12 GB+ or an M4 Max). Speech in English and Chinese; text 30+ languages. Re-check the licence file at your revision.
Streaming speech (and video frames)
Streaming speech
Real-time speech conversation in English and Chinese; text in 30+ languages.
Configurable voices.
Vendor efficiency table: time to first token 0.6 s; decoding 154 tok/s (bf16) and 212 tok/s (int4).
Self-hosted.
Self-hosted.
bf16 about 19 GB, int4 about 11 GB (vendor table). llama.cpp-omni full-duplex: NVIDIA 12 GB+ or Apple M4 Max 24 GB+. PyTorch web demo: 28 GB+.
Apache-2.0 per the model card (older MiniCPM releases had extra registration terms; re-check the repo licence file).
Features
- full-duplex speech streaming
- full-duplex omni (see, listen, speak)
- decides whether to speak at 1 Hz
- stable long speech output (over 1 min)
- quantized builds
Pricing
No public price list.
No licence fee; you pay for GPU time.
n/a (self-hosted)
Free tier: Open weights
Source: huggingface.co
Setup
- GPU with 12 GB+ (int4/llama.cpp) or about 28 GB for the PyTorch web demo.
- Install the pinned Transformers and minicpmo-utils[all] for TTS/streaming.
- Load with trust_remote_code and use the streaming demo, or serve via vLLM/llama.cpp-omni.
Endpoint
Local (your server)
Authentication
n/a
Quick start python
pip install "transformers==4.51.0" accelerate "torch>=2.3.0,<=2.8.0" "torchaudio<=2.8.0" "minicpmo-utils[all]>=1.0.5"
# python
from transformers import AutoModel
model = AutoModel.from_pretrained("openbmb/MiniCPM-o-4_5", trust_remote_code=True, torch_dtype="auto").eval().cuda()
# or serve (text path) with vLLM:
# vllm serve openbmb/MiniCPM-o-4_5 --trust-remote-code --max-num-batched-tokens 2048 --port 8000
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
trust_remote_code
Loading runs model code from the repo; pin a revision and review it before production.
Pinned dependency versions
Requires transformers 4.51.0 and torch <= 2.8.0; conflicts with newer stacks are likely. Use a dedicated environment or container.
Bilingual speech only
Speech conversation is English and Chinese even though text covers 30+ languages.
Licence history
Earlier MiniCPM models used a custom licence with commercial registration; the 4.5 card says Apache-2.0, but confirm the LICENSE file at the revision you ship.
1 Hz speak decision
Full-duplex turn decisions happen once per second, which can feel slower than Moshi-class models.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Half-duplex speech streaming marked under development
- Speech conversation is bilingual (EN/ZH) only
Models and products
| Name | Status |
|---|---|
| openbmb/MiniCPM-o-4_5 | GA |
| openbmb/MiniCPM-o-2_6 | GA |
Docs and sources
Docs
Sources used
Real-world end-to-end voice latency; whether vLLM serving includes speech output.