GA StepFun

StepAudio Realtime

StepFun's end-to-end voice models (StepAudio 2.5 Realtime, Step-Audio 2) on an OpenAI-Realtime-style WebSocket, with voice cloning and paralinguistic cues like laughs and sighs. Suits Chinese-first companion and role-play apps.

Est. per minuten/a
1 high-severity warning

Overview

Best for: Expressive Chinese-language companion or character voice apps.

At a glance

Audio in $/1M tok$1.5
Audio out $/1M tok$10
Native S2SYes
Own LLMNo
CloningYes
WebRTCNo
WebSocketYes
Phone / SIPNo
Open weightsNo

Prices are stepaudio-2.5-realtime per 1M tokens (cache-hit input $0.30); audio tokens per second not published. Platform concurrency tiers (5 at V0) may not apply to realtime. Docs are Chinese-first.

Audio in

pcm16 (sample rate not stated on the model page)

Audio out

pcm16

Languages

Not listed; docs and examples are Chinese.

Voices

Preset voices such as linjiajiejie; voice cloning via uploaded reference audio returns a custom voice id.

Latency

Not published.

Regions

China platform (platform.stepfun.com, CNY) and a USD-priced platform at platform.stepfun.ai; realtime availability on the .ai platform not confirmed.

Compliance

Not stated.

Features

  • server VAD
  • streaming audio deltas
  • voice cloning
  • persona instructions
  • paralinguistic output (laughs, sighs)

Pricing

WhatPriceUnitNotes
stepaudio-2.5-realtime$1.50 in (cache miss) / $0.30 in (cache hit) / $10.00 outper 1M tokensChina site: 10 / 2 / 70 CNY.
step-audio-2$1.43 / $0.29 / $10.00per 1M tokens
step-1o-audio$3.57 / $0.71 / $8.57per 1M tokens
How the per-minute estimate was worked out

Audio tokens per second are not documented; measure usage on a test call.

Audio token rate

Not published, so per-minute cost cannot be derived from docs.

Free tier: Not confirmed.

Source: platform.stepfun.ai

Setup

  1. Create an account on the StepFun open platform and an API key.
  2. Connect to /v1/realtime with ?model=stepaudio-2.5-realtime and Authorization: Bearer <key>.
  3. Send session.update (instructions, voice, pcm16 formats, server_vad), stream audio, play response.audio.delta.

Endpoint

wss://api.stepfun.com/v1/realtime?model=stepaudio-2.5-realtime

Authentication

Authorization: Bearer <STEPFUN_API_KEY>

Quick start javascript

// npm i ws   (China platform host shown; check the console for the international host)
import WebSocket from "ws";

const ws = new WebSocket("wss://api.stepfun.com/v1/realtime?model=stepaudio-2.5-realtime", {
  headers: { Authorization: `Bearer ${process.env.STEPFUN_API_KEY}` },
});

ws.on("open", () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      modalities: ["text", "audio"],
      instructions: "You are a warm, concise assistant.",
      voice: "linjiajiejie",
      input_audio_format: "pcm16",
      output_audio_format: "pcm16",
      turn_detection: { type: "server_vad", prefix_padding_ms: 500 },
    },
  }));
});

export function sendPcm(buf) {
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: buf.toString("base64") }));
}

ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  if (ev.type === "response.audio.delta") play(Buffer.from(ev.delta, "base64"));
});
function play(pcm) {}

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Cost per minute is unknowable from docs

Billing is per token but the audio tokens-per-second rate is not published. Run a metered test call before quoting customers.

Model string inconsistency

Docs use stepaudio-2.5-realtime, press used step-2.5-realtime, and third-party lists vary. Check the console model list before hardcoding.

Chinese-first docs

The detailed realtime docs are Chinese-only; language support, function calling and session limits are not documented.

International availability unclear

A USD price list exists on platform.stepfun.ai but it does not say realtime models are available there; the documented endpoint is api.stepfun.com.

Concurrency tied to top-up

Platform concurrency scales with cumulative spend (5 concurrent at the lowest tier), if those tiers apply to realtime.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Account rate tiers on the open platform range from V0 (under $15 top-up: 5 concurrency, 100 RPM) to V4 ($1,500+: 130 concurrency, 2,600 RPM); not stated whether these apply to realtime
  • Session and context limits not documented on the model page

Models and products

NameStatusNotes
stepaudio-2.5-realtimeGAReleased May 2026 per press; roleplay-focused RLHF. A news article used the string step-2.5-realtime, docs use stepaudio-2.5-realtime.
step-audio-2GAPrevious generation.
step-audio-2-miniGAListed in the realtime guide as a model value; not on the price table. Open weights exist (see open-model entry).
step-1o-audioGAOlder model, still priced.

Docs and sources

Docs

Sources used

Not fully verified

Audio token rate, sample rates, languages, function calling, international endpoint, free tier.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?