Open models

Weights you can run on your own GPU. Hardware needs and licence terms are listed for each one.

Columns 10/16
Basics
Hardware
Licence
Features
Performance
Warnings
Compare 0
Must have:

Showing 30 of 30. Click any column header to sort.

faster-whisper (engine for DIY streaming)SYSTRAN (open source)
Speech-to-textGA--YesMITYes-990
Kyutai STT (Delayed Streams Modeling)Kyutai
Speech-to-textGA1--CC-BY-4.0Yes50020
Moonshine Voice (Moonshine v2 streaming)Moonshine AI (Useful Sensors)
Speech-to-textGA--YesMITYes--0
NVIDIA Nemotron ASR Streaming / Parakeet (Riva, NIM, NeMo)NVIDIA
Speech-to-textGA0.6-NoOpenMDW-1.1Yes-400
whisper_streaming and SimulStreaming (UFAL)Charles University UFAL (open source)
Speech-to-textGA---MITYes-990
WhisperLive (Collabora)Collabora (open source)
Speech-to-textGA--YesMITYes-990
WhisperLiveKitOpen source (QuentinFuxa)
Speech-to-textGA---Apache-2.0Yes-990
Chatterbox (Turbo, Nano, Multilingual V3)Resemble AI
Text-to-speechGA0.35-YesMITYes103231
F5-TTSSWivid (academic)
Text-to-speechGA---CC-BY-NC-4.0No-21
Fun-Audio-Chat-8BFunAudioLLM (Alibaba Tongyi Fun team, attribution from repo branding)
Voice-to-voiceGA9.524-Apache-2.0Yes-20
GLM-4-Voice-9BZhipu AI (zai-org)
Voice-to-voiceGA9--Custom--20
Kimi-Audio-7B-InstructMoonshot AI
Voice-to-voiceGA10--MITYes-20
Kokoro-82Mhexgrad (community)
Text-to-speechGA0.082-YesApache-2.0Yes-80
Kyutai Pocket TTSKyutai
Text-to-speechGA--YesCC-BY-4.0Yes20060
Kyutai TTS (Delayed Streams Modeling)Kyutai
Text-to-speechGA1.6-NoCC-BY-4.0Yes-20
LFM2.5-Audio-1.5BLiquid AI
Voice-to-voiceGA1.5-YesCustomNo-11
Microsoft VibeVoice (Realtime-0.5B)Microsoft
Text-to-speechBeta0.5--MITYes20011
MiniCPM-o 4.5OpenBMB (ModelBest / Tsinghua)
Voice-to-voiceGA911-Apache-2.0Yes60020
MoshiKyutai
Voice-to-voiceGA7.724NoCC-BY-4.0Yes20011
Nari Labs Dia / Dia2Nari Labs
Text-to-speechGA2--Apache-2.0Yes-10
NeuTTS Air / NanoNeuphonic
Text-to-speechGA--YesApache-2.0Yes-10
Orpheus TTS (Canopy Labs)Canopy Labs
Text-to-speechGA3-NoApache-2.0-20010
PersonaPlex-7BNVIDIA
Voice-to-voiceGA824NoCustomYes-10
Qwen3-Omni-30B-A3B (open weights)Alibaba Qwen
Voice-to-voiceGA3579NoApache-2.0Yes-102
Qwen3-TTSAlibaba Qwen
Text-to-speechGA1.7--Apache-2.0Yes97100
Sesame CSM-1BSesame
Text-to-speechGA1--Apache-2.0Yes-10
Step-Audio 2 miniStepFun
Voice-to-voiceGA8--Apache-2.0Yes-20
UnmuteKyutai
Voice-to-voiceGA-16NoCC-BY-4.0Yes75020
VoxCPM2OpenBMB
Text-to-speechGA2--Apache-2.0Yes-300
Coqui XTTS-v2Coqui (defunct); maintained fork by Idiap
Text-to-speechDeprecated--NoCustomNo200171

A dash means the vendor does not say. Latency is the vendor's own claim, not our measurement. Per-minute figures are estimates from list prices; each API page explains the basis, because vendors bill by tokens, characters, connection time or flat minutes.

Category warnings

These apply to most APIs in this category.

Open-weight licences are not all permissive

Apache-2.0/MIT: Qwen3-Omni, MiniCPM-o 4.5, Step-Audio 2 mini, Kimi-Audio, Fun-Audio-Chat, Sesame CSM, Unmute code. CC-BY-4.0 (attribution): Moshi and Kyutai STT/TTS. Custom: NVIDIA Open Model License (PersonaPlex), GLM-4 licence (GLM-4-Voice), LFM Open License with a $10M revenue cap (Liquid LFM2.5-Audio).

Open models rarely ship a production realtime server

Only Moshi/PersonaPlex (moshi.server), Unmute (docker compose) and MiniCPM-o (llama.cpp-omni full duplex) come close. Qwen3-Omni's vLLM path did not output speech at release. Budget engineering time for VAD, streaming, barge-in, auth and scaling.

Free tiers and open weights are not always commercial

Gradium Free is explicitly non-commercial; F5-TTS weights are CC-BY-NC, XTTS-v2 uses the non-commercial Coqui licence (no one left to sell a commercial one), Mistral Voxtral TTS weights and Fish Audio S2 Pro weights are non-commercial. Apache-2.0/MIT options: Kokoro, Chatterbox, Qwen3-TTS, VoxCPM2, Dia, Sesame CSM, NeuTTS Air.