Open models
Weights you can run on your own GPU. Hardware needs and licence terms are listed for each one.
Columns 10/16
Showing 30 of 30. Click any column header to sort.
faster-whisper (engine for DIY streaming)SYSTRAN (open source) |
Speech-to-text | GA | - | - | Yes | MIT | Yes | - | No | - | 99 | 0 | 3 | SYSTRAN (open source) | medium | 2026-10-10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Speech-to-text | GA | 1 | - | - | CC-BY-4.0 | Yes | - | Yes | 500 | 2 | 0 | 4 | Kyutai | medium | 2026-10-10 | |
Moonshine Voice (Moonshine v2 streaming)Moonshine AI (Useful Sensors) |
Speech-to-text | GA | - | - | Yes | MIT | Yes | - | Yes | - | - | 0 | 3 | Moonshine AI (Useful Sensors) | medium | 2026-10-10 |
| Speech-to-text | GA | 0.6 | - | No | OpenMDW-1.1 | Yes | - | Yes | - | 40 | 0 | 4 | NVIDIA | medium | 2026-10-10 | |
whisper_streaming and SimulStreaming (UFAL)Charles University UFAL (open source) |
Speech-to-text | GA | - | - | - | MIT | Yes | - | Yes | - | 99 | 0 | 3 | Charles University UFAL (open source) | medium | 2026-10-10 |
WhisperLive (Collabora)Collabora (open source) |
Speech-to-text | GA | - | - | Yes | MIT | Yes | - | Yes | - | 99 | 0 | 3 | Collabora (open source) | medium | 2026-10-10 |
WhisperLiveKitOpen source (QuentinFuxa) |
Speech-to-text | GA | - | - | - | Apache-2.0 | Yes | - | Yes | - | 99 | 0 | 3 | Open source (QuentinFuxa) | medium | 2026-10-10 |
Chatterbox (Turbo, Nano, Multilingual V3)Resemble AI |
Text-to-speech | GA | 0.35 | - | Yes | MIT | Yes | - | No | 103 | 23 | 1 | 4 | Resemble AI | high | 2026-10-10 |
F5-TTSSWivid (academic) |
Text-to-speech | GA | - | - | - | CC-BY-NC-4.0 | No | - | No | - | 2 | 1 | 3 | SWivid (academic) | high | 2026-10-10 |
Fun-Audio-Chat-8BFunAudioLLM (Alibaba Tongyi Fun team, attribution from repo branding) |
Voice-to-voice | GA | 9.5 | 24 | - | Apache-2.0 | Yes | - | - | - | 2 | 0 | 4 | FunAudioLLM (Alibaba Tongyi Fun team, attribution from repo branding) | medium | 2026-10-10 |
GLM-4-Voice-9BZhipu AI (zai-org) |
Voice-to-voice | GA | 9 | - | - | Custom | - | No | Yes | - | 2 | 0 | 4 | Zhipu AI (zai-org) | medium | 2026-10-10 |
Kimi-Audio-7B-InstructMoonshot AI |
Voice-to-voice | GA | 10 | - | - | MIT | Yes | No | Yes | - | 2 | 0 | 4 | Moonshot AI | medium | 2026-10-10 |
Kokoro-82Mhexgrad (community) |
Text-to-speech | GA | 0.082 | - | Yes | Apache-2.0 | Yes | - | No | - | 8 | 0 | 4 | hexgrad (community) | high | 2026-10-10 |
Kyutai Pocket TTSKyutai |
Text-to-speech | GA | - | - | Yes | CC-BY-4.0 | Yes | - | Yes | 200 | 6 | 0 | 3 | Kyutai | medium | 2026-10-10 |
| Text-to-speech | GA | 1.6 | - | No | CC-BY-4.0 | Yes | - | Yes | - | 2 | 0 | 3 | Kyutai | medium | 2026-10-10 | |
LFM2.5-Audio-1.5BLiquid AI |
Voice-to-voice | GA | 1.5 | - | Yes | Custom | No | No | - | - | 1 | 1 | 4 | Liquid AI | high | 2026-10-10 |
Microsoft VibeVoice (Realtime-0.5B)Microsoft |
Text-to-speech | Beta | 0.5 | - | - | MIT | Yes | - | Yes | 200 | 1 | 1 | 3 | Microsoft | high | 2026-10-10 |
MiniCPM-o 4.5OpenBMB (ModelBest / Tsinghua) |
Voice-to-voice | GA | 9 | 11 | - | Apache-2.0 | Yes | Yes | Yes | 600 | 2 | 0 | 5 | OpenBMB (ModelBest / Tsinghua) | high | 2026-10-10 |
MoshiKyutai |
Voice-to-voice | GA | 7.7 | 24 | No | CC-BY-4.0 | Yes | Yes | Yes | 200 | 1 | 1 | 6 | Kyutai | high | 2026-10-10 |
Nari Labs Dia / Dia2Nari Labs |
Text-to-speech | GA | 2 | - | - | Apache-2.0 | Yes | - | Yes | - | 1 | 0 | 3 | Nari Labs | medium | 2026-10-10 |
NeuTTS Air / NanoNeuphonic |
Text-to-speech | GA | - | - | Yes | Apache-2.0 | Yes | - | - | - | 1 | 0 | 3 | Neuphonic | medium | 2026-10-10 |
Orpheus TTS (Canopy Labs)Canopy Labs |
Text-to-speech | GA | 3 | - | No | Apache-2.0 | - | - | Yes | 200 | 1 | 0 | 4 | Canopy Labs | high | 2026-10-10 |
PersonaPlex-7BNVIDIA |
Voice-to-voice | GA | 8 | 24 | No | Custom | Yes | Yes | Yes | - | 1 | 0 | 5 | NVIDIA | medium | 2026-10-10 |
Qwen3-Omni-30B-A3B (open weights)Alibaba Qwen |
Voice-to-voice | GA | 35 | 79 | No | Apache-2.0 | Yes | No | - | - | 10 | 2 | 5 | Alibaba Qwen | high | 2026-10-10 |
Qwen3-TTSAlibaba Qwen |
Text-to-speech | GA | 1.7 | - | - | Apache-2.0 | Yes | - | Yes | 97 | 10 | 0 | 3 | Alibaba Qwen | high | 2026-10-10 |
Sesame CSM-1BSesame |
Text-to-speech | GA | 1 | - | - | Apache-2.0 | Yes | - | No | - | 1 | 0 | 3 | Sesame | medium | 2026-10-10 |
Step-Audio 2 miniStepFun |
Voice-to-voice | GA | 8 | - | - | Apache-2.0 | Yes | No | - | - | 2 | 0 | 4 | StepFun | medium | 2026-10-10 |
UnmuteKyutai |
Voice-to-voice | GA | - | 16 | No | CC-BY-4.0 | Yes | No | Yes | 750 | 2 | 0 | 5 | Kyutai | high | 2026-10-10 |
VoxCPM2OpenBMB |
Text-to-speech | GA | 2 | - | - | Apache-2.0 | Yes | - | Yes | - | 30 | 0 | 3 | OpenBMB | medium | 2026-10-10 |
Coqui XTTS-v2Coqui (defunct); maintained fork by Idiap |
Text-to-speech | Deprecated | - | - | No | Custom | No | - | Yes | 200 | 17 | 1 | 3 | Coqui (defunct); maintained fork by Idiap | high | 2026-10-10 |
A dash means the vendor does not say. Latency is the vendor's own claim, not our measurement. Per-minute figures are estimates from list prices; each API page explains the basis, because vendors bill by tokens, characters, connection time or flat minutes.
Category warnings
These apply to most APIs in this category.
Open-weight licences are not all permissive
Apache-2.0/MIT: Qwen3-Omni, MiniCPM-o 4.5, Step-Audio 2 mini, Kimi-Audio, Fun-Audio-Chat, Sesame CSM, Unmute code. CC-BY-4.0 (attribution): Moshi and Kyutai STT/TTS. Custom: NVIDIA Open Model License (PersonaPlex), GLM-4 licence (GLM-4-Voice), LFM Open License with a $10M revenue cap (Liquid LFM2.5-Audio).
Open models rarely ship a production realtime server
Only Moshi/PersonaPlex (moshi.server), Unmute (docker compose) and MiniCPM-o (llama.cpp-omni full duplex) come close. Qwen3-Omni's vLLM path did not output speech at release. Budget engineering time for VAD, streaming, barge-in, auth and scaling.
Free tiers and open weights are not always commercial
Gradium Free is explicitly non-commercial; F5-TTS weights are CC-BY-NC, XTTS-v2 uses the non-commercial Coqui licence (no one left to sell a commercial one), Mistral Voxtral TTS weights and Fish Audio S2 Pro weights are non-commercial. Apache-2.0/MIT options: Kokoro, Chatterbox, Qwen3-TTS, VoxCPM2, Dia, Sesame CSM, NeuTTS Air.