GA Canopy Labs

Orpheus TTS (Canopy Labs)

3B-parameter Llama-based expressive TTS with emotion tags, streaming via vLLM. Same model family Groq now hosts.

Est. per minuten/a

Overview

Best for: Expressive English with laughs/sighs when you control GPUs.

At a glance

TypeText-to-speech
Params B3
CPU okNo
LicenceApache-2.0
StreamingYes
Latency ms200
Languages1

About 200 ms streaming latency (about 100 ms with input streaming). Apache-2.0 per card but built on Llama 3.2 3B, so Llama licence terms may apply. Non-English checkpoints are research previews.

Audio in

Text with tags like <laugh>, <sigh>

Audio out

24 kHz 16-bit PCM chunks

Languages

English (prod); others research-grade

Voices

Preset voices (e.g. tara); zero-shot cloning via pretrained model

Latency

Vendor README: ~200 ms streaming latency, reducible to ~100 ms with input streaming.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

GPU recommended (3B LLM backbone); llama.cpp path for CPU exists but is slow.

Licence

Apache-2.0 per model card; built on a Llama 3.2 3B backbone, so check whether Llama licence terms also apply to your use

Features

  • emotion tags
  • output streaming
  • vLLM serving
  • llama.cpp no-GPU option

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. pip install orpheus-speech (pulls vLLM).
  2. Accept the gated model on Hugging Face and log in.
  3. Call generate_speech and write chunks as they arrive.

Endpoint

Local

Authentication

Hugging Face token to download

Quick start python

# pip install orpheus-speech   (GPU + vLLM)
import wave
from orpheus_tts import OrpheusModel

model = OrpheusModel(model_name="canopylabs/orpheus-tts-0.1-finetune-prod")
chunks = model.generate_speech(prompt="Hey there! <laugh> Long time no see.", voice="tara")

with wave.open("out.wav", "wb") as wf:
    wf.setnchannels(1); wf.setsampwidth(2); wf.setframerate(24000)
    for chunk in chunks:          # PCM bytes as they are generated
        wf.writeframes(chunk)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Llama-derived weights

The card says Apache-2.0, but the model is built on Llama-3b. Have counsel confirm whether Meta's Llama licence obligations flow through for commercial use.

Heavy for a TTS

A 3B LLM per stream needs real GPU memory; concurrency per GPU is far lower than Kokoro-class models.

Multilingual models are research previews

Non-English checkpoints are labelled research release.

Hosted alternative is limited

Groq's hosted Orpheus caps requests at 200 characters and WAV only.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Needs an LLM inference stack (vLLM) for real-time

Models and products

NameStatusNotes
canopylabs/orpheus-3b-0.1-ftReleased 2025-03English finetune (gated, auto-approval).
Multilingual research releaseResearch preview (2025-04)7 pretrain/finetune pairs (fr, de, ko, hi, zh, es/it).

Docs and sources

Docs

Sources used

Not fully verified

Production model name in snippet (from README) may change.

Similar open models APIs

Spotted a wrong price or a dead link?