Qwen3-Omni-30B-A3B (open weights)
Open-weight version of Qwen's omni model: hears speech, sees images/video and talks back, under Apache-2.0. Powerful and multilingual, but heavy to run and the standard vLLM server does not produce speech yet.
Overview
Best for: Self-hosted multilingual multimodal assistants where licence freedom (Apache-2.0) matters and data-centre GPUs are available.
At a glance
30B-A3B MoE, about 35B total and 3B active. 10 speech-output languages; 19 speech input, 119 text. 78.85 GB BF16 is the vendor theoretical minimum (with 15 s video). vLLM served text output only at release.
Speech in 19 languages (18 listed)
Speech in 10 languages
Text 119 languages; speech input about 19; speech output 10 (EN, ZH, FR, DE, RU, IT, ES, PT, JA, KO).
Ethan (default), Chelsie, Aiden.
Not given as a number on the model card.
Self-hosted.
Self-hosted.
Vendor 'theoretical minimum' BF16 memory: 78.85 GB (15 s video) to 144.81 GB (120 s video) for Instruct; think 2x A100/H100 80 GB or one H200 class card. Audio-only use needs less but no official figure.
Apache-2.0
Features
- audio, image and video input
- speech output with 3 voices
- thinking variant
Pricing
No public price list.
No licence fee; you pay for GPU time. Hosted equivalents are on Alibaba Model Studio (see the Qwen-Omni Realtime entry).
n/a (self-hosted)
Free tier: Open weights
Source: huggingface.co
Setup
- Provision about 80 GB+ of GPU memory for BF16.
- Install Transformers from source plus qwen-omni-utils (and flash-attn).
- Load Qwen3OmniMoeForConditionalGeneration and call generate with speaker set; build your own streaming loop for realtime use.
Endpoint
None (library)
Authentication
n/a
Quick start python
pip install git+https://github.com/huggingface/transformers accelerate
pip install -U qwen-omni-utils flash-attn --no-build-isolation
# python
from transformers import Qwen3OmniMoeForConditionalGeneration, Qwen3OmniMoeProcessor
m = "Qwen/Qwen3-Omni-30B-A3B-Instruct"
model = Qwen3OmniMoeForConditionalGeneration.from_pretrained(m, dtype="auto", device_map="auto", attn_implementation="flash_attention_2")
proc = Qwen3OmniMoeProcessor.from_pretrained(m)
# build inputs with qwen_omni_utils.process_mm_info, then:
# text_ids, audio = model.generate(**inputs, speaker="Ethan")
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
No speech output from vLLM
The recommended vLLM path served only the text thinker at release; speech output needs the Transformers path (slow) or newer tooling. Check current vLLM support before planning.
Big GPU bill
Roughly 80 GB+ VRAM in BF16 per the vendor table; not a single consumer GPU model.
Not full duplex
It is a turn-based omni model; you must build VAD, streaming and barge-in yourself.
Hosted versions are newer
Alibaba's hosted realtime models (qwen3.5/3.8-omni) are not open-weight; the open release is the 2025 Qwen3-Omni.
Only 3 voices
Voice choice is limited to Ethan, Chelsie and Aiden.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- vLLM serving supports only the thinker (no audio output) at the time of the card
- No turn-key realtime/duplex server in the release
Models and products
| Name | Status |
|---|---|
| Qwen/Qwen3-Omni-30B-A3B-Instruct | GA |
| Qwen/Qwen3-Omni-30B-A3B-Thinking | GA |
| Qwen/Qwen2.5-Omni-7B / 3B | GA |
Docs and sources
Docs
Sources used
Current vLLM / vLLM-Omni audio-output support; audio-only VRAM; exact processor/generate call shape in the snippet.