StepAudio Realtime
StepFun's end-to-end voice models (StepAudio 2.5 Realtime, Step-Audio 2) on an OpenAI-Realtime-style WebSocket, with voice cloning and paralinguistic cues like laughs and sighs. Suits Chinese-first companion and role-play apps.
Overview
Best for: Expressive Chinese-language companion or character voice apps.
At a glance
Prices are stepaudio-2.5-realtime per 1M tokens (cache-hit input $0.30); audio tokens per second not published. Platform concurrency tiers (5 at V0) may not apply to realtime. Docs are Chinese-first.
pcm16 (sample rate not stated on the model page)
pcm16
Not listed; docs and examples are Chinese.
Preset voices such as linjiajiejie; voice cloning via uploaded reference audio returns a custom voice id.
Not published.
China platform (platform.stepfun.com, CNY) and a USD-priced platform at platform.stepfun.ai; realtime availability on the .ai platform not confirmed.
Not stated.
Features
- server VAD
- streaming audio deltas
- voice cloning
- persona instructions
- paralinguistic output (laughs, sighs)
Pricing
| What | Price | Unit |
|---|---|---|
| stepaudio-2.5-realtime | $1.50 in (cache miss) / $0.30 in (cache hit) / $10.00 out | per 1M tokens |
| step-audio-2 | $1.43 / $0.29 / $10.00 | per 1M tokens |
| step-1o-audio | $3.57 / $0.71 / $8.57 | per 1M tokens |
Audio tokens per second are not documented; measure usage on a test call.
Not published, so per-minute cost cannot be derived from docs.
Free tier: Not confirmed.
Source: platform.stepfun.ai
Setup
- Create an account on the StepFun open platform and an API key.
- Connect to /v1/realtime with ?model=stepaudio-2.5-realtime and Authorization: Bearer <key>.
- Send session.update (instructions, voice, pcm16 formats, server_vad), stream audio, play response.audio.delta.
Endpoint
wss://api.stepfun.com/v1/realtime?model=stepaudio-2.5-realtime
Authentication
Authorization: Bearer <STEPFUN_API_KEY>
Quick start javascript
// npm i ws (China platform host shown; check the console for the international host)
import WebSocket from "ws";
const ws = new WebSocket("wss://api.stepfun.com/v1/realtime?model=stepaudio-2.5-realtime", {
headers: { Authorization: `Bearer ${process.env.STEPFUN_API_KEY}` },
});
ws.on("open", () => {
ws.send(JSON.stringify({
type: "session.update",
session: {
modalities: ["text", "audio"],
instructions: "You are a warm, concise assistant.",
voice: "linjiajiejie",
input_audio_format: "pcm16",
output_audio_format: "pcm16",
turn_detection: { type: "server_vad", prefix_padding_ms: 500 },
},
}));
});
export function sendPcm(buf) {
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: buf.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.audio.delta") play(Buffer.from(ev.delta, "base64"));
});
function play(pcm) {}
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Cost per minute is unknowable from docs
Billing is per token but the audio tokens-per-second rate is not published. Run a metered test call before quoting customers.
Model string inconsistency
Docs use stepaudio-2.5-realtime, press used step-2.5-realtime, and third-party lists vary. Check the console model list before hardcoding.
Chinese-first docs
The detailed realtime docs are Chinese-only; language support, function calling and session limits are not documented.
International availability unclear
A USD price list exists on platform.stepfun.ai but it does not say realtime models are available there; the documented endpoint is api.stepfun.com.
Concurrency tied to top-up
Platform concurrency scales with cumulative spend (5 concurrent at the lowest tier), if those tiers apply to realtime.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Account rate tiers on the open platform range from V0 (under $15 top-up: 5 concurrency, 100 RPM) to V4 ($1,500+: 130 concurrency, 2,600 RPM); not stated whether these apply to realtime
- Session and context limits not documented on the model page
Models and products
| Name | Status |
|---|---|
| stepaudio-2.5-realtime | GA |
| step-audio-2 | GA |
| step-audio-2-mini | GA |
| step-1o-audio | GA |
Docs and sources
Docs
Sources used
- platform.stepfun.com/docs/zh/guides/models/stepaudio-2.5-realtime
- platform.stepfun.ai/docs/en/guides/pricing/details
- marktechpost.com/2026/05/24/stepfun-releases-stepaudio-2-5-realtime-an-end-to-e...
Audio token rate, sample rates, languages, function calling, international endpoint, free tier.