GA SYSTRAN (open source)

faster-whisper (engine for DIY streaming)

CTranslate2 re-implementation of Whisper, 4x-class faster than openai-whisper. Not a streaming server itself; it is the decoding engine most self-hosted live tools (WhisperLive, WhisperLiveKit) build on.

Est. per minute$0

Overview

Best for: Building your own Whisper-based live pipeline or powering WhisperLive/WhisperLiveKit.

At a glance

TypeSpeech-to-text
CPU okYes
LicenceMIT
CommercialYes
StreamingNo
Languages99

Batch decoding library, no streaming API by itself. CPU int8 workable for tiny/base/small. Languages follow Whisper (about 99).

Audio in

Files or numpy arrays (16 kHz).

Audio out

n/a

Languages

Whisper languages.

Latency

Batch engine; latency depends on chunking strategy you build.

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

NVIDIA GPU for large models in real time; CPU int8 workable for tiny/base/small.

Licence

MIT (code); Whisper weights MIT.

Features

  • built-in Silero VAD filter
  • word timestamps
  • int8/float16 quantization
  • batched inference pipeline

Pricing

WhatPriceUnitNotes
Software$0Compute only.
How the per-minute estimate was worked out

Self-host compute only.

Free tier: Open source.

Source: github.com

Setup

  1. pip install faster-whisper
  2. Load a model on CUDA or CPU (int8).
  3. Buffer live audio into short windows, transcribe each, and merge (or use WhisperLive / WhisperLiveKit).

Endpoint

n/a (library)

Authentication

none

Quick start python

from faster_whisper import WhisperModel  # pip install faster-whisper

model = WhisperModel("large-v3-turbo", device="cuda", compute_type="float16")
# Transcribe one buffered window of live audio (e.g. the last few seconds)
segments, info = model.transcribe("window.wav", vad_filter=True, beam_size=1)
for s in segments:
    print(f"[{s.start:.1f}-{s.end:.1f}] {s.text}")

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Not streaming by itself

faster-whisper only transcribes buffers; partial results, stabilization and endpointing are your job (or a wrapper's).

Hallucination on silence

Whisper can invent text on silent or noisy windows; keep vad_filter on and drop low-probability segments.

CUDA/cuDNN version coupling

GPU use depends on matching CTranslate2, CUDA and cuDNN versions; mismatches fail at load time.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • No streaming API: you must buffer audio and re-transcribe windows (or use a wrapper).

Models and products

NameStatusNotes
Whisper tiny..large-v3, large-v3-turbo, distil variants (CTranslate2 conversions)Active projectLast push 2026-10-10.

Docs and sources

Docs

Sources used

Not fully verified

large-v3-turbo model alias availability in the installed version.

Similar open models APIs

Spotted a wrong price or a dead link?