Google Cloud Text-to-Speech (Chirp 3 HD and Gemini-TTS)
Cloud TTS offers Chirp 3 HD voices with bidirectional gRPC streaming (text in, audio out) and prompt-steerable Gemini-TTS models billed per audio token. Newest Gemini 3.8 Flash TTS models are Preview and only on the Gemini Enterprise API.
Overview
Best for: GCP shops needing many languages and regional endpoints; Gemini-TTS for prompt-styled or multi-speaker output.
At a glance
Figures are Chirp 3 HD: 30 voices, 53 locales, $30/1M with the first 1M chars/month free. Text-in streaming is Chirp 3 HD only and still Preview. Gemini-TTS is token-billed with prompt-based style control. Instant custom voice $60/1M. SSML on legacy voices only.
Text, SSML (legacy voices), natural-language prompt for Gemini-TTS
Streaming: PCM (default), ALAW, MULAW, OGG_OPUS. Batch: LINEAR16, ALAW, MULAW, MP3, OGG_OPUS, PCM.
Chirp 3 HD: 53 locales; Gemini-TTS: see per-model list
30 named Chirp 3 HD / Gemini voices (Achernar ... Zubenelgenubi); Instant Custom Voice cloning available
No numeric claim on the pages checked; Gemini-TTS described as 'very low latency'.
Chirp 3 HD GA in global, us, eu, asia-southeast1, europe-west2, asia-northeast1
Google Cloud data terms and regional endpoints; certifications not re-verified here.
Features
- bidirectional text-in/audio-out streaming (Chirp 3 HD, Preview)
- prompt-based style control (Gemini-TTS)
- multi-speaker dialogue (Gemini-TTS)
- regional endpoints (global, us, eu, asia-southeast1, europe-west2, asia-northeast1 for Chirp 3 HD)
- instant custom voice
Pricing
| What | Price | Unit |
|---|---|---|
| Chirp 3: HD voices | $30 | per 1M characters |
| Instant custom voice | $60 | per 1M characters |
| Gemini 2.5 Flash TTS / 2.5 Flash-Lite Preview TTS | $0.50 in / $10.00 out | per 1M text tokens / per 1M audio tokens |
| Gemini 2.5 Pro TTS | $1.00 in / $20.00 out | per 1M tokens |
| Gemini 3.1 Flash TTS (Preview) | $1.00 in / $20.00 out | per 1M tokens |
| Gemini 3.8 Flash TTS (Preview) | $0.50 in / $9.00 out | per 1M tokens |
| Gemini 3.8 Flash-Lite TTS (Preview) | $0.50 in / $6.00 out | per 1M tokens |
| Neural2 / WaveNet / Standard / Studio (legacy) | $16 / $4 / $4 / $160 | per 1M characters |
Chirp 3 HD: 900 chars x $30/1M = $0.027. Gemini: 60 s x 25 tokens = 1,500 audio tokens/min; 3.8 Flash-Lite $0.009, 3.8 Flash $0.0135 (promo), 2.5 Flash $0.015, 2.5 Pro $0.03 (text input cost negligible).
Free tier: Chirp 3 HD: 1M characters/month; WaveNet/Standard 4M; Gemini-TTS: none
Source: cloud.google.com
Setup
- Enable the Cloud Text-to-Speech API in a GCP project with billing.
- Authenticate with Application Default Credentials (gcloud auth application-default login or a service account).
- pip install google-cloud-texttospeech.
- Call streaming_synthesize with a config message followed by text messages.
Endpoint
texttospeech.googleapis.com (gRPC StreamingSynthesize)
Authentication
Google Cloud IAM / Application Default Credentials (OAuth), not a simple API key for streaming
Quick start python
# pip install --upgrade google-cloud-texttospeech
from google.cloud import texttospeech
client = texttospeech.TextToSpeechClient()
config = texttospeech.StreamingSynthesizeConfig(
voice=texttospeech.VoiceSelectionParams(
name="en-US-Chirp3-HD-Charon", language_code="en-US"))
def requests():
yield texttospeech.StreamingSynthesizeRequest(streaming_config=config)
for text in ["Hello there. ", "How are you ", "today?"]: # e.g. LLM tokens
yield texttospeech.StreamingSynthesizeRequest(
input=texttospeech.StreamingSynthesisInput(text=text))
with open("out.pcm", "wb") as f:
for resp in client.streaming_synthesize(requests()):
f.write(resp.audio_content) # PCM by default for streaming
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Bidi streaming only for Chirp 3 HD and still Preview
The text-in streaming page says StreamingSynthesize is only compatible with Chirp 3 HD voices and is a Pre-GA feature with limited support. Gemini-TTS streams output but you still send the text up front.
Gemini 3.8 TTS is not on the Cloud TTS API
Gemini 3.8 Flash and Flash-Lite TTS are Preview and only available through the Gemini Enterprise API. Their prices double on Jan 1 2027.
Whitespace and SSML tags are billed
Google counts every character including spaces, newlines and SSML tags (except <mark>). Strip markdown and extra whitespace from LLM output before sending.
Token-billed Gemini voices scale with audio length
Gemini-TTS bills 25 audio tokens per second, so slow pacing prompts or long pauses raise cost independent of text length.
gRPC and IAM auth
Streaming uses gRPC with Google credentials; browsers cannot call it directly. Run a server-side relay.
Studio voices are very expensive
Legacy Studio voices are $160/1M characters; do not pick them for agents.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Gemini-TTS: 8,192 input tokens, 16,384 output tokens per request
- StreamingSynthesize: first message must be config only; Preview (Pre-GA terms)
- Quotas per project; see quotas page
Models and products
| Name | Status |
|---|---|
| Chirp 3: HD (e.g. en-US-Chirp3-HD-Charon) | GA |
| gemini-2.5-flash-tts | GA |
| gemini-2.5-pro-tts | GA |
| gemini-2.5-flash-lite-preview-tts | Preview |
| gemini-3.1-flash-tts-preview | Preview |
| Gemini 3.8 Flash TTS / 3.8 Flash-Lite TTS | Preview (Gemini Enterprise API only) |
| Instant Custom Voice (Chirp 3) | GA (restricted) |
Docs and sources
Docs
Sources used
- cloud.google.com/text-to-speech/pricing
- docs.cloud.google.com/text-to-speech/docs/create-audio-text-streaming
- docs.cloud.google.com/text-to-speech/docs/chirp3-hd
- docs.cloud.google.com/text-to-speech/docs/gemini-tts
Latency figures; per-project streaming quotas.