GA Collabora (open source)

WhisperLive (Collabora)

Self-hosted near-live Whisper transcription server with WebSocket clients, VAD, and faster-whisper, TensorRT or OpenVINO backends.

Est. per minute$0

Overview

Best for: Self-hosted multilingual live captions on your own GPU with Whisper accuracy.

At a glance

TypeSpeech-to-text
CPU okYes
LicenceMIT
CommercialYes
StreamingYes
Languages99

Server around any Whisper size; partials are re-decoded windows, not native streaming. CPU workable with tiny/base or OpenVINO. Session length and client count are server settings.

Audio in

Client captures mic or file and streams 16 kHz float audio.

Audio out

n/a

Languages

Whisper's ~99 languages (model-dependent).

Latency

Depends on model size and GPU; Whisper is not natively streaming, so partials are re-decoded windows.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

NVIDIA GPU recommended for small and larger models; CPU possible with tiny/base or OpenVINO.

Licence

MIT

Features

  • VAD
  • multiple concurrent clients (max_clients)
  • max_connection_time
  • browser extension and iOS clients in repo
  • translation

Pricing

WhatPriceUnitNotes
Software$0Compute only.
How the per-minute estimate was worked out

Self-host compute only.

Free tier: Open source.

Source: github.com

Setup

  1. pip install whisper-live
  2. python3 run_server.py --port 9090 --backend faster_whisper
  3. Connect with the Python client (below) or the browser extension.

Endpoint

ws://localhost:9090

Authentication

None by default; put it behind your own auth/TLS proxy.

Quick start python

# server: python3 run_server.py --port 9090 --backend faster_whisper
from whisper_live.client import TranscriptionClient  # pip install whisper-live

client = TranscriptionClient("localhost", 9090, lang="en", model="small", use_vad=True)
client()  # streams the default microphone and prints segments

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

No auth out of the box

The WebSocket server has no authentication; exposing port 9090 publicly lets anyone use your GPU. Put it behind a TLS proxy with auth.

Whisper is not a true streaming model

Partials come from repeatedly decoding a sliding window, which costs more GPU per stream and can rewrite earlier words.

TensorRT backend is fiddly

The project recommends the Docker setup for TensorRT; native builds are version-sensitive.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Server-side max_clients and max_connection_time settings cap concurrency and session length.

Models and products

NameStatusNotes
Any Whisper size via faster_whisper / tensorrt / openvino backendsActive projectLast push 2026-10-07.

Docs and sources

Docs

Sources used

Not fully verified

Client constructor arguments taken from the README pattern; not executed.

Similar open models APIs

Spotted a wrong price or a dead link?