GA Sesame

Sesame CSM-1B

1B conversational speech model that conditions on prior dialogue context. Apache-2.0, in Hugging Face Transformers since 4.52.1. Not a low-latency streaming engine out of the box.

Est. per minuten/a

Overview

Best for: Research into context-aware conversational speech.

At a glance

TypeText-to-speech
Params B1
LicenceApache-2.0
CommercialYes
StreamingNo
Languages1

No official streaming path. Gated download.

Audio in

Text plus optional conversation context (text + audio segments)

Audio out

24 kHz waveform (generator.sample_rate)

Languages

English (other languages weak, per README)

Voices

Speaker IDs and context-based voice continuation

Latency

No streaming or latency claim in the README.

Regions

Wherever you deploy it

Compliance

Your own deployment; no vendor data processing

Hardware

CUDA GPU recommended.

Licence

Apache-2.0

Features

  • context-conditioned conversational prosody
  • Transformers integration

Pricing

WhatPriceUnitNotes
Weights$0You pay for your own compute
How the per-minute estimate was worked out

Self-hosted: cost is your GPU/CPU time, not per character

Free tier: Open weights

Source: huggingface.co

Setup

  1. git clone https://github.com/SesameAILabs/csm and install requirements (CUDA GPU).
  2. Accept the gated model on Hugging Face.
  3. Call load_csm_1b and generate.

Endpoint

Local

Authentication

Hugging Face token

Quick start python

# from the SesameAILabs/csm repo (run inside the cloned repo)
import torchaudio
from generator import load_csm_1b

generator = load_csm_1b(device="cuda")
audio = generator.generate(text="Hello from Sesame.", speaker=0, context=[],
                           max_audio_length_ms=10_000)
torchaudio.save("audio.wav", audio.unsqueeze(0).cpu(), generator.sample_rate)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

The open model is not the demo

The public 1B checkpoint is a base generation model; it does not reproduce Sesame's hosted conversational demo out of the box.

Not built for streaming agents

No official streaming path; latency is high compared with purpose-built realtime TTS.

Gated download

Requires accepting terms on Hugging Face before download.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • No official streaming
  • English focus

Models and products

NameStatusNotes
sesame/csm-1bReleased 2025-03 (gated, auto-approval)Only public Sesame checkpoint.

Docs and sources

Docs

Sources used

Not fully verified

Hardware minimums.

Similar open models APIs

Spotted a wrong price or a dead link?