GA Alibaba Cloud

Alibaba Cloud Model Studio Qwen-Omni-Realtime

Realtime audio and video conversation API for Qwen Omni models over WebSocket or WebRTC, with an OpenAI-style event protocol. By far the cheapest frontier option per minute and strong in Chinese and Asian languages.

Est. per minute$0.0018 - 0.015
2 high-severity warnings

Overview

Best for: Very cost-sensitive voice agents, Chinese and Asian-language markets, and audio-plus-video assistants.

At a glance

Audio in $/1M tok$0.93
Audio out $/1M tok$1.87
Free tierYes
Native S2SYes
ToolsYes
Image inYes
Own LLMNo
CloningYes
Languages36
Context tokens196,608
Max session min120
WebRTCYes
WebSocketYes
Phone / SIPNo
EU dataNo
Open weightsNo

36 speech-output languages; recognition covers 113. Prices are qwen3.8-omni-flash-realtime, Singapore. Free quota is 1M tokens per model for 90 days (Singapore only). Regions are Singapore and Beijing only.

Audio in

PCM 16 kHz (16-bit mono LE); multichannel 2 or 4 channel spatial input supported (double tokens); JPG images about 1 fps for video.

Audio out

PCM 24 kHz.

Languages

Speech recognition 113 languages and dialects; speech generation 36 languages and dialects.

Voices

Multiple built-in voices (default Tina; others include Ethan, longanlingxin) plus voice cloning.

Latency

No numeric vendor claim captured.

Regions

Singapore (ap-southeast-1, International) and China Beijing (cn-beijing); separate API keys per region.

Compliance

Data retention, residency guarantees and certifications for Model Studio realtime were not captured; review Alibaba Cloud International terms before sending regulated data.

Features

  • server_vad, semantic_vad (filters backchannels and noise) or manual turns (WebSocket only)
  • function calling and remote MCP tools (no extra MCP fee)
  • web search (cannot be combined with tool calling)
  • image and video frame input
  • voice cloning
  • multichannel spatial audio input

Pricing

WhatPriceUnitNotes
qwen3.8-omni-flash-realtime text/image/video input$0.23per 1M tokensSingapore (International)
qwen3.8-omni-flash-realtime audio input$0.93per 1M tokens
qwen3.8-omni-flash-realtime text output$0.70per 1M tokensspeech output bills audio AND its matching text
qwen3.8-omni-flash-realtime audio output$1.87per 1M tokens
qwen3.5-omni-plus-realtime$2.10 text in, $16.50 audio in, $12.40 text out, $62.00 audio outper 1M tokensonly audio billed for speech output
qwen3.5-omni-flash-realtime$0.55 text in, $4.50 audio in, $3.30 text out, $17.70 audio outper 1M tokens
Free quota1M tokens per modelSingapore only, valid 90 days from activation or model release
How the per-minute estimate was worked out

Low = qwen3.8-omni-flash-realtime, 1 min user audio (420 tokens x $0.93/1M) + 1 min model audio (750 tokens x $1.87/1M), single turn, excluding the matching output text tokens (small). High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Alibaba documents that each turn re-bills retained history. qwen3.5-omni-plus-realtime: $0.053 low, $0.27 high.

Audio token rate

qwen3.8-omni-flash-realtime and Qwen3.5-Omni-Realtime: input 7 tokens/s (420/min), output 12.5 tokens/s (750/min); qwen3-omni-flash-realtime-2025-09-15: 12.5 tokens/s both ways. Audio under 1 s billed as 1 s.

Free tier: 1M free tokens per model in the Singapore region for 90 days.

Source: alibabacloud.com

Setup

  1. Create an Alibaba Cloud account, activate Model Studio in the Singapore (International) region, and note your workspace id.
  2. Create a Model Studio (DashScope) API key for that region.
  3. Server: open the WebSocket with Authorization: Bearer <key> and ?model=qwen3.8-omni-flash-realtime.
  4. Send session.update (modalities, voice, instructions, turn_detection), stream input_audio_buffer.append, play response.audio.delta.
  5. Browser: use the WebRTC SDP exchange at /api/v1/webrtc/realtime via your server (VAD mode only).

Endpoint

wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime (Beijing: cn-beijing host); WebRTC: https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime

Authentication

Bearer API key on your server; each region needs its own key. No documented browser ephemeral token: relay or proxy the WebRTC SDP exchange through your backend.

Quick start javascript

import WebSocket from "ws"; // server side; keep the DashScope / Model Studio key off the client
const WS_ID = process.env.MODEL_STUDIO_WORKSPACE_ID; // required for qwen3.8-omni-flash-realtime
const url = `wss://${WS_ID}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime`;
const ws = new WebSocket(url, { headers: { Authorization: `Bearer ${process.env.DASHSCOPE_API_KEY}` } });
ws.on("open", () => ws.send(JSON.stringify({
  type: "session.update",
  session: {
    modalities: ["text", "audio"],
    voice: "Tina",
    instructions: "You are a friendly support agent. Keep answers short.",
    turn_detection: { type: "semantic_vad" },
  },
})));
export function sendAudio(pcm16) { // input must be 16 kHz mono PCM16
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "response.audio.delta") playPcm24k(Buffer.from(ev.delta, "base64")); // 24 kHz PCM out
  if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
  if (ev.type === "error") console.error(ev.error);
});

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Context replay drives the bill

Input tokens accumulate: every response re-bills the retained prior audio, images and text as input. Long sessions on 3.5-plus can cost 10x the first minute. Keep sessions short or trim history.

Model ids churn fast

In about a year the recommended id moved from qwen-omni-turbo-realtime to qwen3-omni-flash-realtime to qwen3.5 to qwen3.8. Older ids vanish from docs without a clear retirement notice on the page. Pin dated snapshots where offered and watch the Model Studio deprecation page.

History re-billed per turn

Alibaba explicitly bills retained history again on each turn, so per-minute cost rises over a call; still very low on the 3.8 flash model.

Plus model is not cheap

qwen3.5-omni-plus-realtime audio output is $62/1M, about the same as OpenAI gpt-realtime-2.1. Do not assume all Qwen realtime models are budget options.

Regions and keys are separate

Singapore and Beijing use different hosts, keys and prices; the free quota is Singapore only. Mainland China endpoints carry their own data and regulatory implications.

Protocol differs between model generations

qwen3.8 examples use a nested session config while Qwen3.5 examples use flat input_audio_format/output_audio_format fields; output is response.audio.delta (beta-style naming), not OpenAI GA names.

3.8 bills audio and text for speech output

For qwen3.8-omni-flash-realtime, spoken output is billed as both audio tokens and the matching text tokens; 3.5 models bill only the audio. Spatial (multichannel) input doubles input-audio tokens.

Workspace endpoint and new SDK required

qwen3.8 only works on the workspace-specific host and needs DashScope Python SDK 1.26.5+ or Java 2.22.15+. Old dashscope-intl.aliyuncs.com examples on the web will not work for it.

Region choice is a compliance decision

Singapore and Beijing are separate deployments with separate keys and different prices. Beijing (Chinese mainland) means data processed in China and typically a China-verified account.

Feature conflicts

Web search cannot be combined with tool calling; WebRTC supports only VAD mode; representation_compact cannot change after audio starts.

Tools vs web search

Web search and tool calling are mutually exclusive in a session; MCP tools need a public HTTPS Streamable HTTP server and approval is on by default.

Python SDK sends transcription by default

The Python SDK defaults enable_input_audio_transcription to True; set it explicitly if you do not want transcription events (and any associated cost).

Session hard cap

Sessions end at 120 minutes; build reconnect logic for long-running calls.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Single WebSocket session up to 120 minutes
  • qwen3.8-omni-flash-realtime: up to 196,608 input tokens; retains 100 audio turns, 50 video turns, 600 s of audio and 240 s of video
  • qwen3.5-omni-flash-realtime retains 80 audio turns, 480 s audio, 120 s video
  • Base64 images under 256 KB; max 1080p
  • Concurrency limits on a separate rate-limiting page (not captured)

Models and products

NameStatusNotes
qwen3.8-omni-flash-realtimeGARecommended; WebSocket, WebRTC and AOQ; 196,608 input tokens; workspace id required.
qwen3.5-omni-plus-realtimeGALarger, much more expensive model.
qwen3.5-omni-flash-realtimeGAPrevious flash model, shorter retained history.
qwen3-omni-flash-realtime(-2025-09-15)GAOlder; 12.5 tokens/s for both input and output audio.

Docs and sources

Docs

Sources used

Not fully verified

Concurrency limits, latency, data retention and compliance; full voice list; exact session.update shape for qwen3.8 (examples differ between generations); Beijing prices.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?