Doubao Realtime Voice Model (end-to-end realtime speech)
ByteDance's end-to-end speech-to-speech model behind the Doubao app's voice chat, sold through Volcengine in mainland China. Version 3.0 'Seeduplex' is full-duplex with an OpenAI-Realtime-style JSON protocol; it suits Chinese-language consumer and companion products.
Overview
Best for: Chinese-language consumer voice products, role-play/companion apps and in-car or device assistants targeting mainland China.
At a glance
Prices converted from 80 / 300 CNY per 1M tokens at 7.1 CNY per USD. Free quota exists but size not documented. Mainland China only with real-name verification. Default 60 session starts per minute and 100k TPM per AppID. Connection released after 10 min idle. Older 2.0 models have 12K context.
PCM 16 kHz mono int16 little-endian (Opus also accepted and converted server-side); send 20 ms / 640-byte packets at real-time pace
Ogg Opus by default; PCM 24 kHz mono (32-bit float or s16le) on request
Chinese and English (vendor says other languages are not guaranteed, especially for cloned voices).
O/O2.0: vv, xiaohe, yunzhou, xiaotian Chinese voices; O2.0 adds English voices Tim, Dacey, Stokie. SC/SC2.0: 21 official cloned character voices; custom voice cloning sold separately.
Vendor describes it as low latency; no millisecond figure on the API page.
Chinese mainland (openspeech.bytedance.com). No equivalent speech-to-speech API was found on BytePlus (international).
Not stated for this API in the docs read.
Features
- full-duplex (3.0)
- barge-in
- server VAD, push-to-talk, text input and audio-file input modes
- function calling with parallel calls (3.0)
- system prompt / persona fields
- voice cloning (SC2.0, cloning 2.0 product)
- singing
- web search via extension
- hot words
- conversation history injection and resume (keeps last 20 rounds)
Pricing
| What | Price | Unit |
|---|---|---|
| Input audio | 80 CNY | per 1M tokens |
| Input text | 10 CNY | per 1M tokens |
| Cached input (text or audio) | 5 CNY | per 1M tokens |
| Output audio | 300 CNY | per 1M tokens |
| Output text | 30 CNY | per 1M tokens |
Own estimate at about 7.1 CNY per USD: 30 s user + 30 s agent speech is about 0.015 CNY input + 0.225 CNY output audio plus cached context, about 0.25-0.35 CNY/min; an agent speaking the full minute is about 0.45 CNY of output audio.
Input audio about 6.25 tokens per second; output audio about 25 tokens per second; ratios may change as the model updates (source: Doubao Speech billing page).
Free tier: Free quota exists and can offset cached, uncached and output tokens, but the size was not found in the docs read.
Source: docs.volcengine.com
Setup
- Create a Volcengine account and complete real-name verification (Chinese ID or business licence expected).
- In the Doubao Speech console, enable the end-to-end realtime voice model and create an API key (new console).
- Read the 'access must-read' page: 3.0 uses a new JSON event protocol that differs from the older binary protocol.
- Open the WebSocket with X-Api-Key, send session.create with session.model = 1.2.6.1, then stream 20 ms PCM frames.
- Handle response.output_audio.delta (Ogg Opus by default) and close with session.close.
Endpoint
wss://openspeech.bytedance.com/api/v3/duplex/realtime/dialogue (3.0 full duplex); legacy: wss://openspeech.bytedance.com/api/v3/realtime/dialogue
Authentication
3.0: X-Api-Key header (new console). Legacy: X-Api-App-ID, X-Api-Access-Key, X-Api-Resource-Id: volc.speech.dialog, X-Api-App-Key: PlgvMymc7f3tQnJ6 (fixed value)
Quick start python
# pip install websockets (Doubao Realtime 3.0 "Seeduplex", full-duplex JSON protocol)
import asyncio, base64, json, os, websockets
URL = "wss://openspeech.bytedance.com/api/v3/duplex/realtime/dialogue"
HEADERS = {"X-Api-Key": os.environ["VOLC_SPEECH_API_KEY"]} # new console > API Key management
async def main(pcm_chunks): # 16 kHz mono int16 LE, 20 ms (640-byte) chunks, sent in real time
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
await ws.send(json.dumps({
"type": "session.create",
"session": {"model": "1.2.6.1", # fixed value for the full-duplex version
"instructions": "You are a friendly assistant."},
}))
async def uplink():
for chunk in pcm_chunks:
await ws.send(json.dumps({"type": "input_audio_buffer.append",
"audio": base64.b64encode(chunk).decode()}))
await asyncio.sleep(0.02) # real-time pace: too fast or too slow is an error
async def downlink():
async for msg in ws:
ev = json.loads(msg)
if ev["type"] == "response.output_audio.delta":
pass # base64 Ogg Opus by default; decode and play
elif ev["type"] == "error":
print(ev)
await asyncio.gather(uplink(), downlink())
await ws.send(json.dumps({"type": "session.close"})) # wait for reply before closing
# asyncio.run(main(your_chunks))
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
China-only account and KYC
Sold on Volcengine in mainland China; expect real-name verification (personal ID or business licence) before paid API use, a Chinese-language console and CNY billing. No BytePlus international equivalent was found.
Two incompatible protocols
The 3.0 Seeduplex endpoint uses JSON Realtime-style events; O/SC/2.0 use a custom binary framing on a different path. Code and samples are not interchangeable; migrate event names first, then move asr/tts/dialog config into 'extension'.
Strict real-time pacing
Sending audio faster or slower than real time triggers server errors, and stopping the uplink without a mute event causes timeouts. File-based tests must sleep 20 ms per 20 ms chunk.
Output audio is the cost driver
Output audio is 300 CNY per 1M tokens at 25 tokens per second, so talkative agents cost far more than listeners. Token ratios are documented as subject to change; bill on metered tokens, not estimates.
Low default throughput
60 sessions started per minute and 100k tokens per minute per AppID by default; a busy call centre needs a quota increase arranged with sales in advance.
Chinese-first quality
Docs state only Chinese and English; cloned voices are only reliably good in Chinese.
Postpaid billing lag
Postpaid bills are issued hourly with possible delays of several hours, so keep a balance buffer to avoid suspension.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Default QPM 60 (StartSession / session.create per minute per AppID) and TPM 100,000; raise via sales
- Server releases the connection after 10 minutes with no interaction (error 45000003)
- Uplink audio must keep real-time pace; send input_audio_mute.commit when the mic is muted or the session times out
- Close with session.close and wait for the reply, otherwise error 55000001 ContextCanceled
- O2.0/SC2.0 max context 12K
Models and products
| Name | Status |
|---|---|
| Doubao Realtime 3.0 (Seeduplex), session.model = 1.2.6.1 | GA |
| O2.0 / SC2.0 (half-duplex S2S) | GA |
| O / SC (1.x) | GA |
Docs and sources
Docs
Sources used
- docs.volcengine.com/docs/6561/1594356
- docs.volcengine.com/docs/DoubaoVoice/access-mustread?lang=zh
- docs.volcengine.com/docs/DoubaoVoice/endtoend-realtime-voice-full-duplex-versio...
- docs.volcengine.com/docs/6561/1359370?lang=zh
Free-quota size and resource-pack prices; whether the worked-example rates apply identically to 3.0 Seeduplex and to 2.0; exact session.create payload fields beyond session.model; KYC rules for foreign users.