GA OpenBMB (ModelBest / Tsinghua)

MiniCPM-o 4.5

A 9B open omni model that does full-duplex speech (and video) streaming and runs on a single 12-24 GB GPU or a Mac via llama.cpp. The most practical open full-duplex model for English/Chinese with vision.

Est. per minuten/a

Overview

Best for: On-prem or edge full-duplex voice plus vision assistants on a single mid-range GPU or Mac.

At a glance

TypeVoice-to-voice
Params B9
VRAM GB11
LicenceApache-2.0
CommercialYes
Full duplexYes
StreamingYes
Latency ms600
Languages2

600 ms is vendor time to first token. 11 GB is int4 (bf16 about 19 GB; llama.cpp full duplex needs 12 GB+ or an M4 Max). Speech in English and Chinese; text 30+ languages. Re-check the licence file at your revision.

Audio in

Streaming speech (and video frames)

Audio out

Streaming speech

Languages

Real-time speech conversation in English and Chinese; text in 30+ languages.

Voices

Configurable voices.

Latency

Vendor efficiency table: time to first token 0.6 s; decoding 154 tok/s (bf16) and 212 tok/s (int4).

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

bf16 about 19 GB, int4 about 11 GB (vendor table). llama.cpp-omni full-duplex: NVIDIA 12 GB+ or Apple M4 Max 24 GB+. PyTorch web demo: 28 GB+.

Licence

Apache-2.0 per the model card (older MiniCPM releases had extra registration terms; re-check the repo licence file).

Features

  • full-duplex speech streaming
  • full-duplex omni (see, listen, speak)
  • decides whether to speak at 1 Hz
  • stable long speech output (over 1 min)
  • quantized builds

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. GPU with 12 GB+ (int4/llama.cpp) or about 28 GB for the PyTorch web demo.
  2. Install the pinned Transformers and minicpmo-utils[all] for TTS/streaming.
  3. Load with trust_remote_code and use the streaming demo, or serve via vLLM/llama.cpp-omni.

Endpoint

Local (your server)

Authentication

n/a

Quick start python

pip install "transformers==4.51.0" accelerate "torch>=2.3.0,<=2.8.0" "torchaudio<=2.8.0" "minicpmo-utils[all]>=1.0.5"

# python
from transformers import AutoModel
model = AutoModel.from_pretrained("openbmb/MiniCPM-o-4_5", trust_remote_code=True, torch_dtype="auto").eval().cuda()

# or serve (text path) with vLLM:
# vllm serve openbmb/MiniCPM-o-4_5 --trust-remote-code --max-num-batched-tokens 2048 --port 8000

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

trust_remote_code

Loading runs model code from the repo; pin a revision and review it before production.

Pinned dependency versions

Requires transformers 4.51.0 and torch <= 2.8.0; conflicts with newer stacks are likely. Use a dedicated environment or container.

Bilingual speech only

Speech conversation is English and Chinese even though text covers 30+ languages.

Licence history

Earlier MiniCPM models used a custom licence with commercial registration; the 4.5 card says Apache-2.0, but confirm the LICENSE file at the revision you ship.

1 Hz speak decision

Full-duplex turn decisions happen once per second, which can feel slower than Moshi-class models.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Half-duplex speech streaming marked under development
  • Speech conversation is bilingual (EN/ZH) only

Models and products

NameStatusNotes
openbmb/MiniCPM-o-4_5GA9B total (SigLip2 + Whisper-medium + CosyVoice2 + Qwen3-8B); int4 and GGUF builds available.
openbmb/MiniCPM-o-2_6GAPrevious generation.

Docs and sources

Docs

Sources used

Not fully verified

Real-world end-to-end voice latency; whether vLLM serving includes speech output.

Similar open models APIs

Spotted a wrong price or a dead link?