WhisperLive (Collabora)
Self-hosted near-live Whisper transcription server with WebSocket clients, VAD, and faster-whisper, TensorRT or OpenVINO backends.
Overview
Best for: Self-hosted multilingual live captions on your own GPU with Whisper accuracy.
At a glance
Server around any Whisper size; partials are re-decoded windows, not native streaming. CPU workable with tiny/base or OpenVINO. Session length and client count are server settings.
Client captures mic or file and streams 16 kHz float audio.
n/a
Whisper's ~99 languages (model-dependent).
Depends on model size and GPU; Whisper is not natively streaming, so partials are re-decoded windows.
Self-hosted.
Self-hosted.
NVIDIA GPU recommended for small and larger models; CPU possible with tiny/base or OpenVINO.
MIT
Features
- VAD
- multiple concurrent clients (max_clients)
- max_connection_time
- browser extension and iOS clients in repo
- translation
Pricing
| What | Price | Unit |
|---|---|---|
| Software | $0 |
Self-host compute only.
Free tier: Open source.
Source: github.com
Setup
- pip install whisper-live
- python3 run_server.py --port 9090 --backend faster_whisper
- Connect with the Python client (below) or the browser extension.
Endpoint
ws://localhost:9090
Authentication
None by default; put it behind your own auth/TLS proxy.
Quick start python
# server: python3 run_server.py --port 9090 --backend faster_whisper
from whisper_live.client import TranscriptionClient # pip install whisper-live
client = TranscriptionClient("localhost", 9090, lang="en", model="small", use_vad=True)
client() # streams the default microphone and prints segments
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
No auth out of the box
The WebSocket server has no authentication; exposing port 9090 publicly lets anyone use your GPU. Put it behind a TLS proxy with auth.
Whisper is not a true streaming model
Partials come from repeatedly decoding a sliding window, which costs more GPU per stream and can rewrite earlier words.
TensorRT backend is fiddly
The project recommends the Docker setup for TensorRT; native builds are version-sensitive.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Server-side max_clients and max_connection_time settings cap concurrency and session length.
Models and products
| Name | Status |
|---|---|
| Any Whisper size via faster_whisper / tensorrt / openvino backends | Active project |
Docs and sources
Docs
Sources used
- raw.githubusercontent.com/collabora/WhisperLive/main/README.md
- api.github.com/repos/collabora/WhisperLive
Client constructor arguments taken from the README pattern; not executed.