Deprecated Hume AI

Hume Octave TTS

LLM-based expressive TTS with voice design from text prompts. Hume is sunsetting its TTS and EVI APIs: access ends 2026-11-13 and account data is deleted afterwards.

Est. per minuten/a
1 high-severity warning

Overview

Best for: Nothing new; existing users should migrate (e.g. to Inworld TTS-2 or ElevenLabs v4 for prompt-steered expressiveness).

At a glance

CloningYes
Text stream inYes
EmotionYes
Latency ms100
Languages11
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

API shuts down 2026-11-13 and account data is deleted. Octave 2 (preview): ~100 ms model latency, 11 languages. Pricing not published.

Audio in

Text plus optional acting description

Audio out

MP3, WAV, PCM

Languages

Octave 2: English, Japanese, Korean, Spanish, French, Portuguese, Italian, German, Russian, Hindi, Arabic

Voices

Voice library, voice design, cloning from ~15 s audio

Latency

Vendor: Octave 2 ~100 ms model latency; instant mode first audio ~200 ms.

Regions

Not specified

Compliance

N/A (sunsetting)

Features

  • input streaming (/v0/tts/stream/input)
  • voice design from text prompts
  • voice cloning
  • instant mode

Pricing

WhatPriceUnitNotes
PricingNot published on hume.ai/pricing at time of researchProduct is being shut down
How the per-minute estimate was worked out

Not verifiable; API ends 2026-11-13

Free tier: Not verified

Source: hume.ai

Setup

  1. Do not start new integrations: access ends 2026-11-13.
  2. Existing users: export anything you need and migrate before that date.

Endpoint

wss://api.hume.ai/v0/tts/stream/input (path per docs; host not re-verified)

Authentication

X-Hume-Api-Key header (not re-verified)

Warnings

API shutdown on 2026-11-13

Hume's docs state TTS and EVI access ends November 13, 2026 at 12:01 a.m. EST and account data is permanently deleted after that date. Migrate now.

Data deletion

Cloned and designed voices stored in your Hume account will be deleted; there is no indication they can be exported to another vendor.

Instant mode restrictions

Instant mode needs a predefined voice and num_generations of 1; voice design requests are slower.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • 5,000 characters per utterance
  • 1,000-character descriptions
  • Up to 5 generations per request

Models and products

NameStatusNotes
Octave 1Sunsetting 2026-11-13English and Spanish, ~200 ms (vendor).
Octave 2 (preview)Sunsetting 2026-11-1311 languages, ~100 ms model latency excluding network (vendor).

Docs and sources

Docs

Sources used

Not fully verified

Pricing; exact WebSocket host/auth header.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?