Step-Audio 2 mini
An 8B open end-to-end audio model from StepFun for speech understanding and speech conversation, with paralinguistic understanding and tool-calling claims. A good open base for Chinese/English voice research; the stronger Step-Audio 2 and 2.5 are API-only.
Overview
Best for: Research and prototyping of Chinese/English speech agents with an Apache-2.0 licence.
At a glance
Turn-based release with a Gradio demo; no streaming server. English and Chinese.
Speech audio
Speech audio
English and Chinese tagged; ASR also benchmarked on Cantonese, Japanese, Arabic and Chinese dialects.
Not documented on the card.
Not stated; real-time streaming not stated on the card.
Self-hosted.
Self-hosted.
Not stated; 8B BF16 weights suggest a 24 GB-class GPU (own estimate).
Apache-2.0
Features
- speech-to-speech conversation
- paralinguistic and non-vocal understanding
- tool calling and multimodal RAG (claims; tool benchmark shown for full Step-Audio 2)
Pricing
No public price list.
No licence fee; you pay for GPU time.
n/a (self-hosted)
Free tier: Open weights
Source: huggingface.co
Setup
- Python 3.10+, PyTorch 2.3+ with CUDA.
- Install the listed dependencies, clone the repo, download weights.
- Run the Gradio web demo or examples.py.
Endpoint
Local Gradio demo
Authentication
n/a
Quick start python
git clone https://github.com/stepfun-ai/Step-Audio2 && cd Step-Audio2
pip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml gradio
huggingface-cli download stepfun-ai/Step-Audio-2-mini --local-dir Step-Audio-2-mini
python web_demo.py # or: python examples.py
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Not a realtime server
The release is a turn-based model with a Gradio demo; streaming, VAD and barge-in are up to you.
Mini is weaker than the API model
Benchmarks show the mini below Step-Audio 2 (e.g. paralinguistic 80.00 vs 83.09); tool-calling results are reported for the full model, not mini.
Old pinned Transformers
Requires transformers 4.49.0; isolate the environment.
Chinese/English focus
Speech conversation quality outside EN/ZH is not documented.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- No documented streaming/duplex server
- No VRAM figure
Models and products
| Name | Status |
|---|---|
| stepfun-ai/Step-Audio-2-mini | GA |
| stepfun-ai/Step-Audio-2-mini-Think, -Base | GA |
| stepfun-ai/Step-Audio-R1 / R1.1 | GA |
Docs and sources
Docs
Sources used
- huggingface.co/stepfun-ai/Step-Audio-2-mini
- huggingface.co/api/models?author=stepfun-ai&search=Audio
VRAM needs, GitHub repo URL in the snippet, streaming support.