Orpheus TTS (Canopy Labs)
3B-parameter Llama-based expressive TTS with emotion tags, streaming via vLLM. Same model family Groq now hosts.
Overview
Best for: Expressive English with laughs/sighs when you control GPUs.
At a glance
About 200 ms streaming latency (about 100 ms with input streaming). Apache-2.0 per card but built on Llama 3.2 3B, so Llama licence terms may apply. Non-English checkpoints are research previews.
Text with tags like <laugh>, <sigh>
24 kHz 16-bit PCM chunks
English (prod); others research-grade
Preset voices (e.g. tara); zero-shot cloning via pretrained model
Vendor README: ~200 ms streaming latency, reducible to ~100 ms with input streaming.
Wherever you deploy it
Your own deployment; no vendor data processing
GPU recommended (3B LLM backbone); llama.cpp path for CPU exists but is slow.
Apache-2.0 per model card; built on a Llama 3.2 3B backbone, so check whether Llama licence terms also apply to your use
Features
- emotion tags
- output streaming
- vLLM serving
- llama.cpp no-GPU option
Pricing
| What | Price | Unit |
|---|---|---|
| Weights | $0 |
Self-hosted: cost is your GPU/CPU time, not per character
Free tier: Open weights
Source: huggingface.co
Setup
- pip install orpheus-speech (pulls vLLM).
- Accept the gated model on Hugging Face and log in.
- Call generate_speech and write chunks as they arrive.
Endpoint
Local
Authentication
Hugging Face token to download
Quick start python
# pip install orpheus-speech (GPU + vLLM)
import wave
from orpheus_tts import OrpheusModel
model = OrpheusModel(model_name="canopylabs/orpheus-tts-0.1-finetune-prod")
chunks = model.generate_speech(prompt="Hey there! <laugh> Long time no see.", voice="tara")
with wave.open("out.wav", "wb") as wf:
wf.setnchannels(1); wf.setsampwidth(2); wf.setframerate(24000)
for chunk in chunks: # PCM bytes as they are generated
wf.writeframes(chunk)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Llama-derived weights
The card says Apache-2.0, but the model is built on Llama-3b. Have counsel confirm whether Meta's Llama licence obligations flow through for commercial use.
Heavy for a TTS
A 3B LLM per stream needs real GPU memory; concurrency per GPU is far lower than Kokoro-class models.
Multilingual models are research previews
Non-English checkpoints are labelled research release.
Hosted alternative is limited
Groq's hosted Orpheus caps requests at 200 characters and WAV only.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- Needs an LLM inference stack (vLLM) for real-time
Models and products
| Name | Status |
|---|---|
| canopylabs/orpheus-3b-0.1-ft | Released 2025-03 |
| Multilingual research release | Research preview (2025-04) |
Docs and sources
Docs
Sources used
Production model name in snippet (from README) may change.