Microsoft VibeVoice (Realtime-0.5B)
MIT-licensed research TTS. The long-form VibeVoice-TTS code was pulled in Sept 2025 over misuse; VibeVoice-Realtime-0.5B (Dec 2025) supports streaming text input with ~200 ms first audible latency, English-focused and single-speaker.
Overview
Best for: Experimenting with LLM-token-to-speech streaming on your own GPU.
At a glance
VibeVoice-Realtime-0.5B; 9 more languages are experimental. MIT but framed by Microsoft as a research release (long-form code was withdrawn after misuse). No voice cloning.
Streamed text
Streamed audio
English (9 experimental languages: de, fr, it, ja, ko, nl, pl, pt, es)
Embedded preset speakers only; custom voices by request to Microsoft
README: ~200 ms first audible latency (hardware dependent).
Wherever you deploy it
Your own deployment; no vendor data processing
NVIDIA GPU (CUDA)
MIT; Microsoft frames it as a research framework
Features
- streaming text input
- long-form generation
- demo web server
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- git clone https://github.com/microsoft/VibeVoice && pip install -e .[streamingtts]
- Run python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B
- Use an NVIDIA Deep Learning Container for CUDA.
Endpoint
Local demo server
Authentication
None
Quick start python
# Shell (from the VibeVoice README):
# git clone https://github.com/microsoft/VibeVoice.git && cd VibeVoice
# pip install -e .[streamingtts]
# python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B
# File-based test:
import subprocess
subprocess.run(["python", "demo/realtime_model_inference_from_file.py",
"--model_path", "microsoft/VibeVoice-Realtime-0.5B",
"--txt_path", "demo/text_examples/1p_vibevoice.txt",
"--speaker_name", "Carter"], check=True)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Research release with a misuse history
Microsoft removed the main VibeVoice-TTS code in Sept 2025 after misuse. Treat as research; it could change or be withdrawn again.
No custom voices
Voice prompts are embedded to limit deepfakes; you cannot clone your own voice.
English only in practice
Other languages are explicitly experimental.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Single speaker
- No user voice cloning
Models and products
| Name | Status |
|---|---|
| microsoft/VibeVoice-Realtime-0.5B | Released 2025-12-03 |
| microsoft/VibeVoice-1.5B | Code disabled (2025-09-05) |
Docs and sources
Docs
Sources used
Transport of the demo server (described as real-time service; WebSocket assumed).