GA Google Cloud

Google Cloud Text-to-Speech (Chirp 3 HD and Gemini-TTS)

Cloud TTS offers Chirp 3 HD voices with bidirectional gRPC streaming (text in, audio out) and prompt-steerable Gemini-TTS models billed per audio token. Newest Gemini 3.8 Flash TTS models are Preview and only on the Gemini Enterprise API.

Est. per minute$0.009 - 0.03
2 high-severity warnings

Overview

Best for: GCP shops needing many languages and regional endpoints; Gemini-TTS for prompt-styled or multi-speaker output.

At a glance

$/1M chars$30
Free tierYes
Voices30
CloningYes
Instant cloneYes
Text stream inYes
EmotionYes
SSMLYes
8 kHz phoneYes
WebRTCNo
WebSocketNo
gRPCYes
EU dataYes
Self-hostNo
Open weightsNo

Figures are Chirp 3 HD: 30 voices, 53 locales, $30/1M with the first 1M chars/month free. Text-in streaming is Chirp 3 HD only and still Preview. Gemini-TTS is token-billed with prompt-based style control. Instant custom voice $60/1M. SSML on legacy voices only.

Audio in

Text, SSML (legacy voices), natural-language prompt for Gemini-TTS

Audio out

Streaming: PCM (default), ALAW, MULAW, OGG_OPUS. Batch: LINEAR16, ALAW, MULAW, MP3, OGG_OPUS, PCM.

Languages

Chirp 3 HD: 53 locales; Gemini-TTS: see per-model list

Voices

30 named Chirp 3 HD / Gemini voices (Achernar ... Zubenelgenubi); Instant Custom Voice cloning available

Latency

No numeric claim on the pages checked; Gemini-TTS described as 'very low latency'.

Regions

Chirp 3 HD GA in global, us, eu, asia-southeast1, europe-west2, asia-northeast1

Compliance

Google Cloud data terms and regional endpoints; certifications not re-verified here.

Features

  • bidirectional text-in/audio-out streaming (Chirp 3 HD, Preview)
  • prompt-based style control (Gemini-TTS)
  • multi-speaker dialogue (Gemini-TTS)
  • regional endpoints (global, us, eu, asia-southeast1, europe-west2, asia-northeast1 for Chirp 3 HD)
  • instant custom voice

Pricing

WhatPriceUnitNotes
Chirp 3: HD voices$30per 1M charactersFirst 1M characters/month free
Instant custom voice$60per 1M charactersNo free tier
Gemini 2.5 Flash TTS / 2.5 Flash-Lite Preview TTS$0.50 in / $10.00 outper 1M text tokens / per 1M audio tokensAudio = 25 tokens per second
Gemini 2.5 Pro TTS$1.00 in / $20.00 outper 1M tokens
Gemini 3.1 Flash TTS (Preview)$1.00 in / $20.00 outper 1M tokens
Gemini 3.8 Flash TTS (Preview)$0.50 in / $9.00 outper 1M tokensThrough Dec 31 2026; $1.00 / $18.00 from Jan 1 2027
Gemini 3.8 Flash-Lite TTS (Preview)$0.50 in / $6.00 outper 1M tokensThrough Dec 31 2026; $1.00 / $12.00 from Jan 1 2027
Neural2 / WaveNet / Standard / Studio (legacy)$16 / $4 / $4 / $160per 1M charactersNot low-latency streaming voices
How the per-minute estimate was worked out

Chirp 3 HD: 900 chars x $30/1M = $0.027. Gemini: 60 s x 25 tokens = 1,500 audio tokens/min; 3.8 Flash-Lite $0.009, 3.8 Flash $0.0135 (promo), 2.5 Flash $0.015, 2.5 Pro $0.03 (text input cost negligible).

Free tier: Chirp 3 HD: 1M characters/month; WaveNet/Standard 4M; Gemini-TTS: none

Source: cloud.google.com

Setup

  1. Enable the Cloud Text-to-Speech API in a GCP project with billing.
  2. Authenticate with Application Default Credentials (gcloud auth application-default login or a service account).
  3. pip install google-cloud-texttospeech.
  4. Call streaming_synthesize with a config message followed by text messages.

Endpoint

texttospeech.googleapis.com (gRPC StreamingSynthesize)

Authentication

Google Cloud IAM / Application Default Credentials (OAuth), not a simple API key for streaming

Quick start python

# pip install --upgrade google-cloud-texttospeech
from google.cloud import texttospeech

client = texttospeech.TextToSpeechClient()
config = texttospeech.StreamingSynthesizeConfig(
    voice=texttospeech.VoiceSelectionParams(
        name="en-US-Chirp3-HD-Charon", language_code="en-US"))

def requests():
    yield texttospeech.StreamingSynthesizeRequest(streaming_config=config)
    for text in ["Hello there. ", "How are you ", "today?"]:  # e.g. LLM tokens
        yield texttospeech.StreamingSynthesizeRequest(
            input=texttospeech.StreamingSynthesisInput(text=text))

with open("out.pcm", "wb") as f:
    for resp in client.streaming_synthesize(requests()):
        f.write(resp.audio_content)  # PCM by default for streaming

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Bidi streaming only for Chirp 3 HD and still Preview

The text-in streaming page says StreamingSynthesize is only compatible with Chirp 3 HD voices and is a Pre-GA feature with limited support. Gemini-TTS streams output but you still send the text up front.

Gemini 3.8 TTS is not on the Cloud TTS API

Gemini 3.8 Flash and Flash-Lite TTS are Preview and only available through the Gemini Enterprise API. Their prices double on Jan 1 2027.

Whitespace and SSML tags are billed

Google counts every character including spaces, newlines and SSML tags (except <mark>). Strip markdown and extra whitespace from LLM output before sending.

Token-billed Gemini voices scale with audio length

Gemini-TTS bills 25 audio tokens per second, so slow pacing prompts or long pauses raise cost independent of text length.

gRPC and IAM auth

Streaming uses gRPC with Google credentials; browsers cannot call it directly. Run a server-side relay.

Studio voices are very expensive

Legacy Studio voices are $160/1M characters; do not pick them for agents.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Gemini-TTS: 8,192 input tokens, 16,384 output tokens per request
  • StreamingSynthesize: first message must be config only; Preview (Pre-GA terms)
  • Quotas per project; see quotas page

Models and products

NameStatusNotes
Chirp 3: HD (e.g. en-US-Chirp3-HD-Charon)GA30 voices, 53 locales; the only voice type supported by StreamingSynthesize (bidi streaming, Preview feature).
gemini-2.5-flash-ttsGASingle and multi-speaker; streaming output PCM/ALAW/MULAW/OGG_OPUS.
gemini-2.5-pro-ttsGAHigher control for podcasts/audiobooks.
gemini-2.5-flash-lite-preview-ttsPreviewSingle speaker.
gemini-3.1-flash-tts-previewPreviewInline tags like [laughs], [sigh].
Gemini 3.8 Flash TTS / 3.8 Flash-Lite TTSPreview (Gemini Enterprise API only)Not available through the Cloud TTS API; includes voice design and voice replication.
Instant Custom Voice (Chirp 3)GA (restricted)$60/1M characters.

Docs and sources

Docs

Sources used

Not fully verified

Latency figures; per-project streaming quotas.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?