Beta Microsoft

Microsoft VibeVoice (Realtime-0.5B)

MIT-licensed research TTS. The long-form VibeVoice-TTS code was pulled in Sept 2025 over misuse; VibeVoice-Realtime-0.5B (Dec 2025) supports streaming text input with ~200 ms first audible latency, English-focused and single-speaker.

Est. per minuten/a
1 high-severity warning

Overview

Best for: Experimenting with LLM-token-to-speech streaming on your own GPU.

At a glance

TypeText-to-speech
Params B0.5
LicenceMIT
CommercialYes
StreamingYes
Latency ms200
Languages1

VibeVoice-Realtime-0.5B; 9 more languages are experimental. MIT but framed by Microsoft as a research release (long-form code was withdrawn after misuse). No voice cloning.

Audio in

Streamed text

Audio out

Streamed audio

Languages

English (9 experimental languages: de, fr, it, ja, ko, nl, pl, pt, es)

Voices

Embedded preset speakers only; custom voices by request to Microsoft

Latency

README: ~200 ms first audible latency (hardware dependent).

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

NVIDIA GPU (CUDA)

Licence

MIT; Microsoft frames it as a research framework

Features

  • streaming text input
  • long-form generation
  • demo web server

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. git clone https://github.com/microsoft/VibeVoice && pip install -e .[streamingtts]
  2. Run python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B
  3. Use an NVIDIA Deep Learning Container for CUDA.

Endpoint

Local demo server

Authentication

None

Quick start python

# Shell (from the VibeVoice README):
#   git clone https://github.com/microsoft/VibeVoice.git && cd VibeVoice
#   pip install -e .[streamingtts]
#   python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B
# File-based test:
import subprocess
subprocess.run(["python", "demo/realtime_model_inference_from_file.py",
                "--model_path", "microsoft/VibeVoice-Realtime-0.5B",
                "--txt_path", "demo/text_examples/1p_vibevoice.txt",
                "--speaker_name", "Carter"], check=True)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Research release with a misuse history

Microsoft removed the main VibeVoice-TTS code in Sept 2025 after misuse. Treat as research; it could change or be withdrawn again.

No custom voices

Voice prompts are embedded to limit deepfakes; you cannot clone your own voice.

English only in practice

Other languages are explicitly experimental.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Single speaker
  • No user voice cloning

Models and products

NameStatusNotes
microsoft/VibeVoice-Realtime-0.5BReleased 2025-12-03Streaming text input; embedded voice prompts only.
microsoft/VibeVoice-1.5BCode disabled (2025-09-05)Weights still on HF but TTS code removed from the repo.

Docs and sources

Docs

Sources used

Not fully verified

Transport of the demo server (described as real-time service; WebSocket assumed).

Similar open models APIs

Spotted a wrong price or a dead link?