GA StepFun

Step-Audio 2 mini

An 8B open end-to-end audio model from StepFun for speech understanding and speech conversation, with paralinguistic understanding and tool-calling claims. A good open base for Chinese/English voice research; the stronger Step-Audio 2 and 2.5 are API-only.

Est. per minuten/a

Overview

Best for: Research and prototyping of Chinese/English speech agents with an Apache-2.0 licence.

At a glance

TypeVoice-to-voice
Params B8
LicenceApache-2.0
CommercialYes
Full duplexNo
Languages2

Turn-based release with a Gradio demo; no streaming server. English and Chinese.

Audio in

Speech audio

Audio out

Speech audio

Languages

English and Chinese tagged; ASR also benchmarked on Cantonese, Japanese, Arabic and Chinese dialects.

Voices

Not documented on the card.

Latency

Not stated; real-time streaming not stated on the card.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

Not stated; 8B BF16 weights suggest a 24 GB-class GPU (own estimate).

Licence

Apache-2.0

Features

  • speech-to-speech conversation
  • paralinguistic and non-vocal understanding
  • tool calling and multimodal RAG (claims; tool benchmark shown for full Step-Audio 2)

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. Python 3.10+, PyTorch 2.3+ with CUDA.
  2. Install the listed dependencies, clone the repo, download weights.
  3. Run the Gradio web demo or examples.py.

Endpoint

Local Gradio demo

Authentication

n/a

Quick start python

git clone https://github.com/stepfun-ai/Step-Audio2 && cd Step-Audio2
pip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml gradio
huggingface-cli download stepfun-ai/Step-Audio-2-mini --local-dir Step-Audio-2-mini
python web_demo.py   # or: python examples.py

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Not a realtime server

The release is a turn-based model with a Gradio demo; streaming, VAD and barge-in are up to you.

Mini is weaker than the API model

Benchmarks show the mini below Step-Audio 2 (e.g. paralinguistic 80.00 vs 83.09); tool-calling results are reported for the full model, not mini.

Old pinned Transformers

Requires transformers 4.49.0; isolate the environment.

Chinese/English focus

Speech conversation quality outside EN/ZH is not documented.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • No documented streaming/duplex server
  • No VRAM figure

Models and products

NameStatusNotes
stepfun-ai/Step-Audio-2-miniGA8B, BF16.
stepfun-ai/Step-Audio-2-mini-Think, -BaseGAReasoning and base variants.
stepfun-ai/Step-Audio-R1 / R1.1GALarger (about 33B) audio reasoning models; gated.

Docs and sources

Docs

Sources used

Not fully verified

VRAM needs, GitHub repo URL in the snippet, streaming support.

Similar open models APIs

Spotted a wrong price or a dead link?