faster-whisper (engine for DIY streaming)
CTranslate2 re-implementation of Whisper, 4x-class faster than openai-whisper. Not a streaming server itself; it is the decoding engine most self-hosted live tools (WhisperLive, WhisperLiveKit) build on.
Overview
Best for: Building your own Whisper-based live pipeline or powering WhisperLive/WhisperLiveKit.
At a glance
Batch decoding library, no streaming API by itself. CPU int8 workable for tiny/base/small. Languages follow Whisper (about 99).
Files or numpy arrays (16 kHz).
n/a
Whisper languages.
Batch engine; latency depends on chunking strategy you build.
Self-hosted.
Self-hosted.
NVIDIA GPU for large models in real time; CPU int8 workable for tiny/base/small.
MIT (code); Whisper weights MIT.
Features
- built-in Silero VAD filter
- word timestamps
- int8/float16 quantization
- batched inference pipeline
Pricing
| What | Price | Unit |
|---|---|---|
| Software | $0 |
Self-host compute only.
Free tier: Open source.
Source: github.com
Setup
- pip install faster-whisper
- Load a model on CUDA or CPU (int8).
- Buffer live audio into short windows, transcribe each, and merge (or use WhisperLive / WhisperLiveKit).
Endpoint
n/a (library)
Authentication
none
Quick start python
from faster_whisper import WhisperModel # pip install faster-whisper
model = WhisperModel("large-v3-turbo", device="cuda", compute_type="float16")
# Transcribe one buffered window of live audio (e.g. the last few seconds)
segments, info = model.transcribe("window.wav", vad_filter=True, beam_size=1)
for s in segments:
print(f"[{s.start:.1f}-{s.end:.1f}] {s.text}")
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Not streaming by itself
faster-whisper only transcribes buffers; partial results, stabilization and endpointing are your job (or a wrapper's).
Hallucination on silence
Whisper can invent text on silent or noisy windows; keep vad_filter on and drop low-probability segments.
CUDA/cuDNN version coupling
GPU use depends on matching CTranslate2, CUDA and cuDNN versions; mismatches fail at load time.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- No streaming API: you must buffer audio and re-transcribe windows (or use a wrapper).
Models and products
| Name | Status |
|---|---|
| Whisper tiny..large-v3, large-v3-turbo, distil variants (CTranslate2 conversions) | Active project |
Docs and sources
Docs
Sources used
large-v3-turbo model alias availability in the installed version.