GA Alibaba Qwen

Qwen3-Omni-30B-A3B (open weights)

Open-weight version of Qwen's omni model: hears speech, sees images/video and talks back, under Apache-2.0. Powerful and multilingual, but heavy to run and the standard vLLM server does not produce speech yet.

Est. per minuten/a
2 high-severity warnings

Overview

Best for: Self-hosted multilingual multimodal assistants where licence freedom (Apache-2.0) matters and data-centre GPUs are available.

At a glance

TypeVoice-to-voice
Params B35
VRAM GB79
CPU okNo
LicenceApache-2.0
CommercialYes
Full duplexNo
Languages10

30B-A3B MoE, about 35B total and 3B active. 10 speech-output languages; 19 speech input, 119 text. 78.85 GB BF16 is the vendor theoretical minimum (with 15 s video). vLLM served text output only at release.

Audio in

Speech in 19 languages (18 listed)

Audio out

Speech in 10 languages

Languages

Text 119 languages; speech input about 19; speech output 10 (EN, ZH, FR, DE, RU, IT, ES, PT, JA, KO).

Voices

Ethan (default), Chelsie, Aiden.

Latency

Not given as a number on the model card.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

Vendor 'theoretical minimum' BF16 memory: 78.85 GB (15 s video) to 144.81 GB (120 s video) for Instruct; think 2x A100/H100 80 GB or one H200 class card. Audio-only use needs less but no official figure.

Licence

Apache-2.0

Features

  • audio, image and video input
  • speech output with 3 voices
  • thinking variant

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time. Hosted equivalents are on Alibaba Model Studio (see the Qwen-Omni Realtime entry).

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. Provision about 80 GB+ of GPU memory for BF16.
  2. Install Transformers from source plus qwen-omni-utils (and flash-attn).
  3. Load Qwen3OmniMoeForConditionalGeneration and call generate with speaker set; build your own streaming loop for realtime use.

Endpoint

None (library)

Authentication

n/a

Quick start python

pip install git+https://github.com/huggingface/transformers accelerate
pip install -U qwen-omni-utils flash-attn --no-build-isolation

# python
from transformers import Qwen3OmniMoeForConditionalGeneration, Qwen3OmniMoeProcessor
m = "Qwen/Qwen3-Omni-30B-A3B-Instruct"
model = Qwen3OmniMoeForConditionalGeneration.from_pretrained(m, dtype="auto", device_map="auto", attn_implementation="flash_attention_2")
proc = Qwen3OmniMoeProcessor.from_pretrained(m)
# build inputs with qwen_omni_utils.process_mm_info, then:
# text_ids, audio = model.generate(**inputs, speaker="Ethan")

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No speech output from vLLM

The recommended vLLM path served only the text thinker at release; speech output needs the Transformers path (slow) or newer tooling. Check current vLLM support before planning.

Big GPU bill

Roughly 80 GB+ VRAM in BF16 per the vendor table; not a single consumer GPU model.

Not full duplex

It is a turn-based omni model; you must build VAD, streaming and barge-in yourself.

Hosted versions are newer

Alibaba's hosted realtime models (qwen3.5/3.8-omni) are not open-weight; the open release is the 2025 Qwen3-Omni.

Only 3 voices

Voice choice is limited to Ethan, Chelsie and Aiden.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • vLLM serving supports only the thinker (no audio output) at the time of the card
  • No turn-key realtime/duplex server in the release

Models and products

NameStatusNotes
Qwen/Qwen3-Omni-30B-A3B-InstructGAMoE, about 35B total params, 3B active; speech output via the 'talker'.
Qwen/Qwen3-Omni-30B-A3B-ThinkingGAText output only (reasoning).
Qwen/Qwen2.5-Omni-7B / 3BGAOlder, smaller omni models (Apache-2.0) if 30B is too big.

Docs and sources

Docs

Sources used

Not fully verified

Current vLLM / vLLM-Omni audio-output support; audio-only VRAM; exact processor/generate call shape in the snippet.

Similar open models APIs

Spotted a wrong price or a dead link?