GA ByteDance Volcengine (Doubao Speech)

Doubao Realtime Voice Model (end-to-end realtime speech)

ByteDance's end-to-end speech-to-speech model behind the Doubao app's voice chat, sold through Volcengine in mainland China. Version 3.0 'Seeduplex' is full-duplex with an OpenAI-Realtime-style JSON protocol; it suits Chinese-language consumer and companion products.

Est. per minute$0.03 - 0.07
2 high-severity warnings

Overview

Best for: Chinese-language consumer voice products, role-play/companion apps and in-car or device assistants targeting mainland China.

At a glance

Audio in $/1M tok$11.27
Audio out $/1M tok$42.25
Free tierYes
Native S2SYes
ToolsYes
Own LLMNo
CloningYes
Languages2
WebRTCNo
WebSocketYes
Phone / SIPNo
EU dataNo
Open weightsNo

Prices converted from 80 / 300 CNY per 1M tokens at 7.1 CNY per USD. Free quota exists but size not documented. Mainland China only with real-name verification. Default 60 session starts per minute and 100k TPM per AppID. Connection released after 10 min idle. Older 2.0 models have 12K context.

Audio in

PCM 16 kHz mono int16 little-endian (Opus also accepted and converted server-side); send 20 ms / 640-byte packets at real-time pace

Audio out

Ogg Opus by default; PCM 24 kHz mono (32-bit float or s16le) on request

Languages

Chinese and English (vendor says other languages are not guaranteed, especially for cloned voices).

Voices

O/O2.0: vv, xiaohe, yunzhou, xiaotian Chinese voices; O2.0 adds English voices Tim, Dacey, Stokie. SC/SC2.0: 21 official cloned character voices; custom voice cloning sold separately.

Latency

Vendor describes it as low latency; no millisecond figure on the API page.

Regions

Chinese mainland (openspeech.bytedance.com). No equivalent speech-to-speech API was found on BytePlus (international).

Compliance

Not stated for this API in the docs read.

Features

  • full-duplex (3.0)
  • barge-in
  • server VAD, push-to-talk, text input and audio-file input modes
  • function calling with parallel calls (3.0)
  • system prompt / persona fields
  • voice cloning (SC2.0, cloning 2.0 product)
  • singing
  • web search via extension
  • hot words
  • conversation history injection and resume (keeps last 20 rounds)

Pricing

WhatPriceUnitNotes
Input audio80 CNYper 1M tokensRates taken from the official worked billing example on the pricing page.
Input text10 CNYper 1M tokens
Cached input (text or audio)5 CNYper 1M tokensContext and system prompt hits from earlier turns.
Output audio300 CNYper 1M tokens
Output text30 CNYper 1M tokensIncludes text of the spoken reply and returned ASR text.
How the per-minute estimate was worked out

Own estimate at about 7.1 CNY per USD: 30 s user + 30 s agent speech is about 0.015 CNY input + 0.225 CNY output audio plus cached context, about 0.25-0.35 CNY/min; an agent speaking the full minute is about 0.45 CNY of output audio.

Audio token rate

Input audio about 6.25 tokens per second; output audio about 25 tokens per second; ratios may change as the model updates (source: Doubao Speech billing page).

Free tier: Free quota exists and can offset cached, uncached and output tokens, but the size was not found in the docs read.

Source: docs.volcengine.com

Setup

  1. Create a Volcengine account and complete real-name verification (Chinese ID or business licence expected).
  2. In the Doubao Speech console, enable the end-to-end realtime voice model and create an API key (new console).
  3. Read the 'access must-read' page: 3.0 uses a new JSON event protocol that differs from the older binary protocol.
  4. Open the WebSocket with X-Api-Key, send session.create with session.model = 1.2.6.1, then stream 20 ms PCM frames.
  5. Handle response.output_audio.delta (Ogg Opus by default) and close with session.close.

Endpoint

wss://openspeech.bytedance.com/api/v3/duplex/realtime/dialogue (3.0 full duplex); legacy: wss://openspeech.bytedance.com/api/v3/realtime/dialogue

Authentication

3.0: X-Api-Key header (new console). Legacy: X-Api-App-ID, X-Api-Access-Key, X-Api-Resource-Id: volc.speech.dialog, X-Api-App-Key: PlgvMymc7f3tQnJ6 (fixed value)

Quick start python

# pip install websockets   (Doubao Realtime 3.0 "Seeduplex", full-duplex JSON protocol)
import asyncio, base64, json, os, websockets

URL = "wss://openspeech.bytedance.com/api/v3/duplex/realtime/dialogue"
HEADERS = {"X-Api-Key": os.environ["VOLC_SPEECH_API_KEY"]}  # new console > API Key management

async def main(pcm_chunks):  # 16 kHz mono int16 LE, 20 ms (640-byte) chunks, sent in real time
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        await ws.send(json.dumps({
            "type": "session.create",
            "session": {"model": "1.2.6.1",  # fixed value for the full-duplex version
                        "instructions": "You are a friendly assistant."},
        }))
        async def uplink():
            for chunk in pcm_chunks:
                await ws.send(json.dumps({"type": "input_audio_buffer.append",
                                          "audio": base64.b64encode(chunk).decode()}))
                await asyncio.sleep(0.02)  # real-time pace: too fast or too slow is an error
        async def downlink():
            async for msg in ws:
                ev = json.loads(msg)
                if ev["type"] == "response.output_audio.delta":
                    pass  # base64 Ogg Opus by default; decode and play
                elif ev["type"] == "error":
                    print(ev)
        await asyncio.gather(uplink(), downlink())
        await ws.send(json.dumps({"type": "session.close"}))  # wait for reply before closing

# asyncio.run(main(your_chunks))

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

China-only account and KYC

Sold on Volcengine in mainland China; expect real-name verification (personal ID or business licence) before paid API use, a Chinese-language console and CNY billing. No BytePlus international equivalent was found.

Two incompatible protocols

The 3.0 Seeduplex endpoint uses JSON Realtime-style events; O/SC/2.0 use a custom binary framing on a different path. Code and samples are not interchangeable; migrate event names first, then move asr/tts/dialog config into 'extension'.

Strict real-time pacing

Sending audio faster or slower than real time triggers server errors, and stopping the uplink without a mute event causes timeouts. File-based tests must sleep 20 ms per 20 ms chunk.

Output audio is the cost driver

Output audio is 300 CNY per 1M tokens at 25 tokens per second, so talkative agents cost far more than listeners. Token ratios are documented as subject to change; bill on metered tokens, not estimates.

Low default throughput

60 sessions started per minute and 100k tokens per minute per AppID by default; a busy call centre needs a quota increase arranged with sales in advance.

Chinese-first quality

Docs state only Chinese and English; cloned voices are only reliably good in Chinese.

Postpaid billing lag

Postpaid bills are issued hourly with possible delays of several hours, so keep a balance buffer to avoid suspension.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Default QPM 60 (StartSession / session.create per minute per AppID) and TPM 100,000; raise via sales
  • Server releases the connection after 10 minutes with no interaction (error 45000003)
  • Uplink audio must keep real-time pace; send input_audio_mute.commit when the mic is muted or the session times out
  • Close with session.close and wait for the reply, otherwise error 55000001 ContextCanceled
  • O2.0/SC2.0 max context 12K

Models and products

NameStatusNotes
Doubao Realtime 3.0 (Seeduplex), session.model = 1.2.6.1GAFull-duplex version, JSON text frames, function calling, endpoint /api/v3/duplex/realtime/dialogue.
O2.0 / SC2.0 (half-duplex S2S)GAListed under the 'historical' end-to-end interface docs with a binary protocol. O = Omni route, SC = Strong Character role-play route; 12K max context.
O / SC (1.x)GANo longer iterated; capabilities are converging into the 2.0 versions.

Docs and sources

Docs

Sources used

Not fully verified

Free-quota size and resource-pack prices; whether the worked-example rates apply identically to 3.0 Seeduplex and to 2.0; exact session.create payload fields beyond session.model; KYC rules for foreign users.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?