GA Moonshot AI

Kimi-Audio-7B-Instruct

Moonshot's open audio foundation model (built on Qwen2.5-7B) for speech recognition, audio understanding and end-to-end speech conversation. Moonshot offers no hosted realtime voice API, so this is the only Kimi voice option.

Est. per minuten/a

Overview

Best for: Research on audio understanding plus speech replies with a permissive MIT licence.

At a glance

TypeVoice-to-voice
Params B10
LicenceMIT
CommercialYes
Full duplexNo
StreamingYes
Languages2

Named 7B (Qwen2.5-7B base), about 10B params total. Chunk-wise streaming detokenizer. English and Chinese.

Audio in

Speech audio

Audio out

Speech at 24 kHz (output_type='both' returns text + audio)

Languages

English and Chinese.

Voices

Not documented.

Latency

Chunk-wise streaming detokenizer for low-latency generation; no figure.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

Not stated on the card; about 10B BF16 params suggests 24 GB+ VRAM (own estimate).

Licence

MIT (Qwen2.5-derived code Apache-2.0).

Features

  • ASR
  • audio QA and captioning
  • emotion and sound-event recognition
  • end-to-end speech conversation

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. CUDA GPU (24 GB-class suggested, own estimate).
  2. pip install the Kimi-Audio package or pull the Docker image.
  3. Load KimiAudio with load_detokenizer=True and call generate.

Endpoint

None (library)

Authentication

n/a

Quick start python

pip install git+https://github.com/MoonshotAI/Kimi-Audio.git "transformers<5"
# or: docker pull moonshotai/kimi-audio:v0.1

# python
from kimia_infer.api.kimia import KimiAudio
model = KimiAudio(model_path="moonshotai/Kimi-Audio-7B-Instruct", load_detokenizer=True)
# wav, text = model.generate(messages, output_type="both")

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No realtime serving

Turn-based library; you build streaming, VAD and interruption.

Aging release

Released April 2025 and little updated since; newer open models (MiniCPM-o 4.5, Qwen3-Omni) are generally preferred.

EN/ZH only

Only English and Chinese are tagged.

Moonshot has no voice API

Kimi's hosted API is text/multimodal chat only; do not expect a managed realtime endpoint.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Not a full-duplex realtime server
  • No hosted Moonshot realtime API found

Models and products

NameStatusNotes
moonshotai/Kimi-Audio-7B-InstructGAAbout 10B params BF16; released April 2025.

Docs and sources

Docs

Sources used

Not fully verified

VRAM; exact import path and generate signature in the snippet.

Similar open models APIs

Spotted a wrong price or a dead link?