Alibaba Cloud Model Studio Qwen-Omni-Realtime
Realtime audio and video conversation API for Qwen Omni models over WebSocket or WebRTC, with an OpenAI-style event protocol. By far the cheapest frontier option per minute and strong in Chinese and Asian languages.
Overview
Best for: Very cost-sensitive voice agents, Chinese and Asian-language markets, and audio-plus-video assistants.
At a glance
36 speech-output languages; recognition covers 113. Prices are qwen3.8-omni-flash-realtime, Singapore. Free quota is 1M tokens per model for 90 days (Singapore only). Regions are Singapore and Beijing only.
PCM 16 kHz (16-bit mono LE); multichannel 2 or 4 channel spatial input supported (double tokens); JPG images about 1 fps for video.
PCM 24 kHz.
Speech recognition 113 languages and dialects; speech generation 36 languages and dialects.
Multiple built-in voices (default Tina; others include Ethan, longanlingxin) plus voice cloning.
No numeric vendor claim captured.
Singapore (ap-southeast-1, International) and China Beijing (cn-beijing); separate API keys per region.
Data retention, residency guarantees and certifications for Model Studio realtime were not captured; review Alibaba Cloud International terms before sending regulated data.
Features
- server_vad, semantic_vad (filters backchannels and noise) or manual turns (WebSocket only)
- function calling and remote MCP tools (no extra MCP fee)
- web search (cannot be combined with tool calling)
- image and video frame input
- voice cloning
- multichannel spatial audio input
Pricing
| What | Price | Unit |
|---|---|---|
| qwen3.8-omni-flash-realtime text/image/video input | $0.23 | per 1M tokens |
| qwen3.8-omni-flash-realtime audio input | $0.93 | per 1M tokens |
| qwen3.8-omni-flash-realtime text output | $0.70 | per 1M tokens |
| qwen3.8-omni-flash-realtime audio output | $1.87 | per 1M tokens |
| qwen3.5-omni-plus-realtime | $2.10 text in, $16.50 audio in, $12.40 text out, $62.00 audio out | per 1M tokens |
| qwen3.5-omni-flash-realtime | $0.55 text in, $4.50 audio in, $3.30 text out, $17.70 audio out | per 1M tokens |
| Free quota | 1M tokens per model |
Low = qwen3.8-omni-flash-realtime, 1 min user audio (420 tokens x $0.93/1M) + 1 min model audio (750 tokens x $1.87/1M), single turn, excluding the matching output text tokens (small). High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Alibaba documents that each turn re-bills retained history. qwen3.5-omni-plus-realtime: $0.053 low, $0.27 high.
qwen3.8-omni-flash-realtime and Qwen3.5-Omni-Realtime: input 7 tokens/s (420/min), output 12.5 tokens/s (750/min); qwen3-omni-flash-realtime-2025-09-15: 12.5 tokens/s both ways. Audio under 1 s billed as 1 s.
Free tier: 1M free tokens per model in the Singapore region for 90 days.
Source: alibabacloud.com
Setup
- Create an Alibaba Cloud account, activate Model Studio in the Singapore (International) region, and note your workspace id.
- Create a Model Studio (DashScope) API key for that region.
- Server: open the WebSocket with Authorization: Bearer <key> and ?model=qwen3.8-omni-flash-realtime.
- Send session.update (modalities, voice, instructions, turn_detection), stream input_audio_buffer.append, play response.audio.delta.
- Browser: use the WebRTC SDP exchange at /api/v1/webrtc/realtime via your server (VAD mode only).
Endpoint
wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime (Beijing: cn-beijing host); WebRTC: https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime
Authentication
Bearer API key on your server; each region needs its own key. No documented browser ephemeral token: relay or proxy the WebRTC SDP exchange through your backend.
Quick start javascript
import WebSocket from "ws"; // server side; keep the DashScope / Model Studio key off the client
const WS_ID = process.env.MODEL_STUDIO_WORKSPACE_ID; // required for qwen3.8-omni-flash-realtime
const url = `wss://${WS_ID}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime`;
const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.DASHSCOPE_API_KEY}` } });
ws.on("open", () => ws.send(JSON.stringify({
type: "session.update",
session: {
modalities: ["text", "audio"],
voice: "Tina",
instructions: "You are a friendly support agent. Keep answers short.",
turn_detection: { type: "semantic_vad" },
},
})));
export function sendAudio(pcm16) { // input must be 16 kHz mono PCM16
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.audio.delta") playPcm24k(Buffer.from(ev.delta, "base64")); // 24 kHz PCM out
if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
if (ev.type === "error") console.error(ev.error);
});
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Context replay drives the bill
Input tokens accumulate: every response re-bills the retained prior audio, images and text as input. Long sessions on 3.5-plus can cost 10x the first minute. Keep sessions short or trim history.
Model ids churn fast
In about a year the recommended id moved from qwen-omni-turbo-realtime to qwen3-omni-flash-realtime to qwen3.5 to qwen3.8. Older ids vanish from docs without a clear retirement notice on the page. Pin dated snapshots where offered and watch the Model Studio deprecation page.
History re-billed per turn
Alibaba explicitly bills retained history again on each turn, so per-minute cost rises over a call; still very low on the 3.8 flash model.
Plus model is not cheap
qwen3.5-omni-plus-realtime audio output is $62/1M, about the same as OpenAI gpt-realtime-2.1. Do not assume all Qwen realtime models are budget options.
Regions and keys are separate
Singapore and Beijing use different hosts, keys and prices; the free quota is Singapore only. Mainland China endpoints carry their own data and regulatory implications.
Protocol differs between model generations
qwen3.8 examples use a nested session config while Qwen3.5 examples use flat input_audio_format/output_audio_format fields; output is response.audio.delta (beta-style naming), not OpenAI GA names.
3.8 bills audio and text for speech output
For qwen3.8-omni-flash-realtime, spoken output is billed as both audio tokens and the matching text tokens; 3.5 models bill only the audio. Spatial (multichannel) input doubles input-audio tokens.
Workspace endpoint and new SDK required
qwen3.8 only works on the workspace-specific host and needs DashScope Python SDK 1.26.5+ or Java 2.22.15+. Old dashscope-intl.aliyuncs.com examples on the web will not work for it.
Region choice is a compliance decision
Singapore and Beijing are separate deployments with separate keys and different prices. Beijing (Chinese mainland) means data processed in China and typically a China-verified account.
Feature conflicts
Web search cannot be combined with tool calling; WebRTC supports only VAD mode; representation_compact cannot change after audio starts.
Tools vs web search
Web search and tool calling are mutually exclusive in a session; MCP tools need a public HTTPS Streamable HTTP server and approval is on by default.
Python SDK sends transcription by default
The Python SDK defaults enable_input_audio_transcription to True; set it explicitly if you do not want transcription events (and any associated cost).
Session hard cap
Sessions end at 120 minutes; build reconnect logic for long-running calls.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Single WebSocket session up to 120 minutes
- qwen3.8-omni-flash-realtime: up to 196,608 input tokens; retains 100 audio turns, 50 video turns, 600 s of audio and 240 s of video
- qwen3.5-omni-flash-realtime retains 80 audio turns, 480 s audio, 120 s video
- Base64 images under 256 KB; max 1080p
- Concurrency limits on a separate rate-limiting page (not captured)
Models and products
| Name | Status |
|---|---|
| qwen3.8-omni-flash-realtime | GA |
| qwen3.5-omni-plus-realtime | GA |
| qwen3.5-omni-flash-realtime | GA |
| qwen3-omni-flash-realtime(-2025-09-15) | GA |
Docs and sources
Docs
Sources used
Concurrency limits, latency, data retention and compliance; full voice list; exact session.update shape for qwen3.8 (examples differ between generations); Beijing prices.