Fun-Audio-Chat-8B
An 8B open large audio-language model for low-latency voice chat with speech-to-speech inference, speech function calling and voice empathy, under Apache-2.0. A newer open alternative for Chinese/English voice agents that need tool use.
Overview
Best for: Self-hosted Chinese/English voice agents that need spoken function calling under a permissive licence.
At a glance
Named 8B, about 9.5B params. Example scripts only, no server. Speech function calling. English and Chinese.
Speech (5 Hz shared backbone, 25 Hz refined head)
Speech
English and Chinese.
Not documented on the card.
No millisecond figure; vendor claims the 5 Hz frame rate cuts GPU hours by nearly 50%.
Self-hosted.
Self-hosted.
About 24 GB GPU memory for inference (vendor); 4x80 GB for training.
Apache-2.0
Features
- speech-to-speech
- speech function calling
- spoken QA
- speech instruction following
- voice empathy
Pricing
No public price list.
No licence fee; you pay for GPU time.
n/a (self-hosted)
Free tier: Open weights
Source: huggingface.co
Setup
- GPU with about 24 GB.
- Clone the repo with submodules, install PyTorch 2.8 (cu128) and requirements, download weights.
- Run the speech-to-speech example script.
Endpoint
None (scripts)
Authentication
n/a
Quick start python
git clone --recurse-submodules https://github.com/FunAudioLLM/Fun-Audio-Chat && cd Fun-Audio-Chat
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
hf download FunAudioLLM/Fun-Audio-Chat-8B --local-dir ./pretrained_models/Fun-Audio-Chat-8B
export PYTHONPATH=`pwd`
python examples/infer_s2s.py
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Example scripts, not a server
You must build streaming, VAD and an API around it.
Strict environment
Python 3.12, PyTorch 2.8.0 and ffmpeg are required; the card's generic Transformers snippet does not match the repo.
EN/ZH only
Only English and Chinese.
Young project
Low download counts and limited community tooling so far.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Full-duplex not stated
- No serving stack beyond example scripts
Models and products
| Name | Status |
|---|---|
| FunAudioLLM/Fun-Audio-Chat-8B | GA |
Docs and sources
Docs
Sources used
Vendor attribution to Alibaba Tongyi (inferred from branding); latency; duplex behaviour.