Google Cloud Speech-to-Text V2 streaming (Chirp 3)
gRPC StreamingRecognize on Speech-to-Text V2 with the chirp_3 model in us/eu multi-regions. Strong language coverage, but short stream limit and no streaming diarization on Chirp 3.
Overview
Best for: GCP-native apps that need wide language coverage and IAM/VPC-SC controls, with short utterances (commands, IVR).
At a glance
Chirp 3 lists about 111 locales (29 GA, 82 Preview), not a language count. $0.016/min V2 standard up to 500K min/month, cheaper at volume. About 5 minutes of audio per stream. $300 is the general Google Cloud trial credit. HIPAA via Google Cloud BAA, not re-verified per model. Chirp 3 diarization is batch only; interim results not confirmed for chirp_3.
ExplicitDecodingConfig (LINEAR16, MULAW, ALAW and others) with sample rate and channel count, or auto-decoding for containers. Audio must be sent at roughly real-time pace.
n/a
Chirp 3 table lists about 111 locales (my count: 29 GA, 82 Preview).
No vendor figure captured.
chirp_3: us and eu multi-regions (GA); more planned. Use the regional API endpoint (e.g. us-speech.googleapis.com) matching the recognizer location.
Google Cloud compliance programs (HIPAA BAA via Google Cloud, etc.) apply per Google's covered-services list; data logging opt-in affects V1 pricing. Not re-verified per model.
Features
- streaming interim results (check per model)
- utterance-level timestamps (streaming only)
- speech adaptation / phrase sets
- language-agnostic auto-detect
- voice activity events
Pricing
| What | Price | Unit |
|---|---|---|
| V2 Standard recognition (incl. chirp) | $0.016 | per minute, 0-500K min/month |
| V2 Standard, 500K-1M | $0.010 | per minute |
| V2 Standard, 1M-2M | $0.008 | per minute |
| V2 Standard, 2M+ | $0.004 | per minute |
| V2 Dynamic batch (not streaming) | $0.003 | per minute |
| V1 API | $0.016 with data logging / $0.024 without | per minute after 60 free min/month |
V2 tiered price; most teams pay $0.016/min until 500K minutes per month. Multi-channel audio is billed per channel.
Free tier: V1 includes 60 free minutes/month; new Google Cloud accounts get $300 trial credit (Google Cloud free program).
Source: cloud.google.com
Setup
- Enable the Speech-to-Text API on a Google Cloud project and authenticate with a service account (ADC).
- pip install google-cloud-speech.
- Create a SpeechClient with api_endpoint set to the regional host (us-speech.googleapis.com for location us).
- Send a first StreamingRecognizeRequest with recognizer projects/<id>/locations/us/recognizers/_ and streaming_config, then audio-only requests.
- Restart the stream before ~5 minutes.
Endpoint
gRPC us-speech.googleapis.com:443 / eu-speech.googleapis.com:443 (Speech.StreamingRecognize, V2)
Authentication
Google Cloud IAM (service account / ADC OAuth tokens). No browser-direct option; proxy through your backend.
Quick start python
from google.api_core.client_options import ClientOptions
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech as cs # pip install google-cloud-speech
PROJECT = "my-project"
client = SpeechClient(client_options=ClientOptions(api_endpoint="us-speech.googleapis.com"))
cfg = cs.RecognitionConfig(
explicit_decoding_config=cs.ExplicitDecodingConfig(
encoding=cs.ExplicitDecodingConfig.AudioEncoding.LINEAR16,
sample_rate_hertz=16000, audio_channel_count=1),
language_codes=["en-US"], model="chirp_3")
scfg = cs.StreamingRecognitionConfig(
config=cfg, streaming_features=cs.StreamingRecognitionFeatures(interim_results=True))
def requests():
yield cs.StreamingRecognizeRequest(
recognizer=f"projects/{PROJECT}/locations/us/recognizers/_", streaming_config=scfg)
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(3200): # 100 ms; send at real-time pace for live audio
yield cs.StreamingRecognizeRequest(audio=chunk)
for resp in client.streaming_recognize(requests=requests()):
for r in resp.results:
if r.alternatives:
print("FINAL" if r.is_final else "partial", r.alternatives[0].transcript)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
About 5 minutes per stream
Streaming requests are limited to roughly 5 minutes of audio. Long calls need you to open a new stream before the limit and stitch results, handling words cut at the boundary.
No streaming diarization on Chirp 3
Chirp 3 speaker diarization is only available in BatchRecognize. Word-level timestamps are also not supported in streaming.
Region-locked endpoint
chirp_3 is only in the us and eu multi-regions; the recognizer location and the API endpoint host must match or calls fail.
Expensive at low volume
$0.016/min until 500K minutes a month is 3x Deepgram Nova-3 PAYG. Discounts only kick in at large volumes.
Many Chirp 3 locales are Preview
Of the ~111 listed locales most are Preview, which carries no SLA; check your language's status before committing.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Streaming requests limited to about 5 minutes of audio per stream; reconnect and stitch for longer audio.
- 10 MB limit applies to StreamingRecognize request messages.
- Phrase sets: 5,000 phrases / 100,000 characters per request (general quotas page); Chirp 3 adaptation dictionary up to 1,000 phrases.
Models and products
| Name | Status |
|---|---|
| chirp_3 | GA (us and eu multi-regions) |
| chirp_2 | Older generation |
| telephony / long / short / latest_* | Legacy standard models |
Docs and sources
Docs
Sources used
- cloud.google.com/speech-to-text/pricing
- docs.cloud.google.com/speech-to-text/v2/docs/chirp_3-model
- docs.cloud.google.com/speech-to-text/quotas
Whether chirp_3 returns interim results in streaming, chirp_2 status, and any V2 free minutes. The V2 pricing table was read from page markup.