GA FunAudioLLM (Alibaba Tongyi Fun team, attribution from repo branding)

Fun-Audio-Chat-8B

An 8B open large audio-language model for low-latency voice chat with speech-to-speech inference, speech function calling and voice empathy, under Apache-2.0. A newer open alternative for Chinese/English voice agents that need tool use.

Est. per minuten/a

Overview

Best for: Self-hosted Chinese/English voice agents that need spoken function calling under a permissive licence.

At a glance

TypeVoice-to-voice
Params B9.5
VRAM GB24
LicenceApache-2.0
CommercialYes
Languages2

Named 8B, about 9.5B params. Example scripts only, no server. Speech function calling. English and Chinese.

Audio in

Speech (5 Hz shared backbone, 25 Hz refined head)

Audio out

Speech

Languages

English and Chinese.

Voices

Not documented on the card.

Latency

No millisecond figure; vendor claims the 5 Hz frame rate cuts GPU hours by nearly 50%.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

About 24 GB GPU memory for inference (vendor); 4x80 GB for training.

Licence

Apache-2.0

Features

  • speech-to-speech
  • speech function calling
  • spoken QA
  • speech instruction following
  • voice empathy

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. GPU with about 24 GB.
  2. Clone the repo with submodules, install PyTorch 2.8 (cu128) and requirements, download weights.
  3. Run the speech-to-speech example script.

Endpoint

None (scripts)

Authentication

n/a

Quick start python

git clone --recurse-submodules https://github.com/FunAudioLLM/Fun-Audio-Chat && cd Fun-Audio-Chat
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
hf download FunAudioLLM/Fun-Audio-Chat-8B --local-dir ./pretrained_models/Fun-Audio-Chat-8B
export PYTHONPATH=`pwd`
python examples/infer_s2s.py

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Example scripts, not a server

You must build streaming, VAD and an API around it.

Strict environment

Python 3.12, PyTorch 2.8.0 and ffmpeg are required; the card's generic Transformers snippet does not match the repo.

EN/ZH only

Only English and Chinese.

Young project

Low download counts and limited community tooling so far.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Full-duplex not stated
  • No serving stack beyond example scripts

Models and products

NameStatusNotes
FunAudioLLM/Fun-Audio-Chat-8BGAAbout 9.5B params; December 2025.

Docs and sources

Docs

Sources used

Not fully verified

Vendor attribution to Alibaba Tongyi (inferred from branding); latency; duplex behaviour.

Similar open models APIs

Spotted a wrong price or a dead link?