GA MiniMax

MiniMax Speech (international)

High-quality multilingual TTS (40 language_boost options) with a bidirectional WebSocket that buffers LLM text and splits on punctuation. Among the most expensive per character.

Est. per minute$0.054 - 0.09
1 high-severity warning

Overview

Best for: Expressive multilingual and Chinese-language voices, cloning and voice design where cost is secondary.

At a glance

$/1M chars$60
CloningYes
Instant cloneYes
Text stream inYes
EmotionYes
8 kHz phoneYes
Languages40
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

$60/1M is speech-2.8-turbo; hd is $100/1M. 40 language_boost values. Rapid voice cloning $1.5 per voice. Limits are RPM by subscription, not concurrency. pcmu output at 8 kHz supported.

Audio in

Text with pause markers <#x#>

Audio out

mp3 (default), pcm, flac, wav, pcmu_raw, pcmu_wav, opus; 8000-44100 Hz (default 32000); mono or stereo

Languages

40 language_boost values plus auto

Voices

System voices; rapid voice cloning ($1.5/voice) and voice design ($3/voice); voice mixing up to 4

Latency

No millisecond figure on the bidi docs.

Regions

International host api.minimax.io (China platform uses api.minimaxi.com with separate keys)

Compliance

Not verified.

Features

  • input streaming (t2a_v2_bidi)
  • server-side sentence segmentation
  • task_cancel for barge-in
  • voice cloning
  • voice design
  • emotion control
  • voice mixing

Pricing

WhatPriceUnitNotes
speech-2.8-hd$100per 1M charactersSync and async
speech-2.8-turbo$60per 1M characters
Rapid voice cloning$1.5per voiceCharged on first use
Voice design$3per voiceCharged on first use
Audio subscription$5 to $999/monthmonthly100K to 20M audio points; RPM 10 to 800
How the per-minute estimate was worked out

900 chars/min; turbo $60/M vs hd $100/M pay-as-you-go

Free tier: Not verified

Source: platform.minimax.io

Setup

  1. Create an account on platform.minimax.io (international) and an API key.
  2. Connect to wss://api.minimax.io/ws/v1/t2a_v2_bidi.
  3. Wait for connected_success, send task_start (model, voice_setting, audio_setting), wait for task_started.
  4. Send task_continue per text chunk, then task_finish; decode audio from data.audio.

Endpoint

wss://api.minimax.io/ws/v1/t2a_v2_bidi

Authentication

Authorization: Bearer <MINIMAX_API_KEY> (assumed from the standard T2A API; not shown on the bidi page excerpt)

Quick start python

# pip install websockets
import asyncio, json, os, websockets

URL = "wss://api.minimax.io/ws/v1/t2a_v2_bidi"
HDR = {"Authorization": f"Bearer {os.environ['MINIMAX_API_KEY']}"}

async def main():
    async with websockets.connect(URL, additional_headers=HDR) as ws:
        await ws.recv()  # connected_success
        await ws.send(json.dumps({"event": "task_start", "model": "speech-2.8-turbo",
            "voice_setting": {"voice_id": "YOUR_VOICE_ID", "speed": 1},
            "audio_setting": {"format": "pcm", "sample_rate": 24000}}))
        await ws.recv()  # task_started
        for t in ["Hello there. ", "This is streamed ", "from an LLM."]:
            await ws.send(json.dumps({"event": "task_continue", "text": t}))
        await ws.send(json.dumps({"event": "task_finish"}))
        with open("out_24k.pcm", "wb") as f:
            async for raw in ws:
                m = json.loads(raw)
                audio = (m.get("data") or {}).get("audio")
                if audio:
                    f.write(bytes.fromhex(audio))  # T2A returns hex-encoded audio
                if m.get("event") == "task_finished":
                    break

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Expensive per character

At $60-100 per 1M characters MiniMax costs roughly 4-7x Deepgram, Inworld or Google Chirp for the same text. Model your volume first.

Short unpunctuated text waits

The bidi server synthesizes immediately only at sentence-final punctuation; short fragments without punctuation wait for a backstop window. Send task_flush at end of turn.

You must send keepalive pings

The server sends no pings and closes after ~120 s of inactivity (error 2201).

China vs international keys

Keys from minimaxi.com (China) and minimax.io (international) are separate platforms; use the host matching your key.

Audio may be hex, not base64

The standard T2A API returns hex-encoded audio; verify encoding on the bidi stream before decoding.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • task_continue text under 10,000 characters
  • ~120 s idle timeout; server sends no pings, client must
  • One synthesis session per connection
  • Subscription RPM: Starter 10 ... Business 800

Models and products

NameStatusNotes
speech-2.8-hdGA$100/M characters.
speech-2.8-turboGA$60/M characters.
speech-2.6-hd / speech-2.6-turboLegacySame prices as 2.8 equivalents.
speech-02-hd / speech-02-turbo / speech-01-*Legacy

Docs and sources

Docs

Sources used

Not fully verified

Auth header on bidi endpoint; hex vs base64 audio on bidi; free credits.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?