GA Google Cloud

Google Cloud Speech-to-Text V2 streaming (Chirp 3)

gRPC StreamingRecognize on Speech-to-Text V2 with the chirp_3 model in us/eu multi-regions. Strong language coverage, but short stream limit and no streaming diarization on Chirp 3.

Est. per minute$0.004 - 0.016
2 high-severity warnings

Overview

Best for: GCP-native apps that need wide language coverage and IAM/VPC-SC controls, with short utterances (commands, IVR).

At a glance

$/hour$0.96
Free tierYes
Free credit $$300
Live speakersNo
KeytermsYes
8 kHz phoneYes
Max session min5
WebRTCNo
WebSocketNo
gRPCYes
HIPAAYes
EU dataYes
Self-hostNo
Open weightsNo

Chirp 3 lists about 111 locales (29 GA, 82 Preview), not a language count. $0.016/min V2 standard up to 500K min/month, cheaper at volume. About 5 minutes of audio per stream. $300 is the general Google Cloud trial credit. HIPAA via Google Cloud BAA, not re-verified per model. Chirp 3 diarization is batch only; interim results not confirmed for chirp_3.

Audio in

ExplicitDecodingConfig (LINEAR16, MULAW, ALAW and others) with sample rate and channel count, or auto-decoding for containers. Audio must be sent at roughly real-time pace.

Audio out

n/a

Languages

Chirp 3 table lists about 111 locales (my count: 29 GA, 82 Preview).

Latency

No vendor figure captured.

Regions

chirp_3: us and eu multi-regions (GA); more planned. Use the regional API endpoint (e.g. us-speech.googleapis.com) matching the recognizer location.

Compliance

Google Cloud compliance programs (HIPAA BAA via Google Cloud, etc.) apply per Google's covered-services list; data logging opt-in affects V1 pricing. Not re-verified per model.

Features

  • streaming interim results (check per model)
  • utterance-level timestamps (streaming only)
  • speech adaptation / phrase sets
  • language-agnostic auto-detect
  • voice activity events

Pricing

WhatPriceUnitNotes
V2 Standard recognition (incl. chirp)$0.016per minute, 0-500K min/monthSame SKU for streaming and sync. 1-second rounding per request.
V2 Standard, 500K-1M$0.010per minute
V2 Standard, 1M-2M$0.008per minute
V2 Standard, 2M+$0.004per minute
V2 Dynamic batch (not streaming)$0.003per minuteDiscounted low-urgency batch.
V1 API$0.016 with data logging / $0.024 withoutper minute after 60 free min/monthV1 only. Medical models $0.078/min.
How the per-minute estimate was worked out

V2 tiered price; most teams pay $0.016/min until 500K minutes per month. Multi-channel audio is billed per channel.

Free tier: V1 includes 60 free minutes/month; new Google Cloud accounts get $300 trial credit (Google Cloud free program).

Source: cloud.google.com

Setup

  1. Enable the Speech-to-Text API on a Google Cloud project and authenticate with a service account (ADC).
  2. pip install google-cloud-speech.
  3. Create a SpeechClient with api_endpoint set to the regional host (us-speech.googleapis.com for location us).
  4. Send a first StreamingRecognizeRequest with recognizer projects/<id>/locations/us/recognizers/_ and streaming_config, then audio-only requests.
  5. Restart the stream before ~5 minutes.

Endpoint

gRPC us-speech.googleapis.com:443 / eu-speech.googleapis.com:443 (Speech.StreamingRecognize, V2)

Authentication

Google Cloud IAM (service account / ADC OAuth tokens). No browser-direct option; proxy through your backend.

Quick start python

from google.api_core.client_options import ClientOptions
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech as cs  # pip install google-cloud-speech

PROJECT = "my-project"
client = SpeechClient(client_options=ClientOptions(api_endpoint="us-speech.googleapis.com"))
cfg = cs.RecognitionConfig(
    explicit_decoding_config=cs.ExplicitDecodingConfig(
        encoding=cs.ExplicitDecodingConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=16000, audio_channel_count=1),
    language_codes=["en-US"], model="chirp_3")
scfg = cs.StreamingRecognitionConfig(
    config=cfg, streaming_features=cs.StreamingRecognitionFeatures(interim_results=True))

def requests():
    yield cs.StreamingRecognizeRequest(
        recognizer=f"projects/{PROJECT}/locations/us/recognizers/_", streaming_config=scfg)
    with open("audio_16k_mono.raw", "rb") as f:
        while chunk := f.read(3200):  # 100 ms; send at real-time pace for live audio
            yield cs.StreamingRecognizeRequest(audio=chunk)

for resp in client.streaming_recognize(requests=requests()):
    for r in resp.results:
        if r.alternatives:
            print("FINAL" if r.is_final else "partial", r.alternatives[0].transcript)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

About 5 minutes per stream

Streaming requests are limited to roughly 5 minutes of audio. Long calls need you to open a new stream before the limit and stitch results, handling words cut at the boundary.

No streaming diarization on Chirp 3

Chirp 3 speaker diarization is only available in BatchRecognize. Word-level timestamps are also not supported in streaming.

Region-locked endpoint

chirp_3 is only in the us and eu multi-regions; the recognizer location and the API endpoint host must match or calls fail.

Expensive at low volume

$0.016/min until 500K minutes a month is 3x Deepgram Nova-3 PAYG. Discounts only kick in at large volumes.

Many Chirp 3 locales are Preview

Of the ~111 listed locales most are Preview, which carries no SLA; check your language's status before committing.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Streaming requests limited to about 5 minutes of audio per stream; reconnect and stitch for longer audio.
  • 10 MB limit applies to StreamingRecognize request messages.
  • Phrase sets: 5,000 phrases / 100,000 characters per request (general quotas page); Chirp 3 adaptation dictionary up to 1,000 phrases.

Models and products

NameStatusNotes
chirp_3GA (us and eu multi-regions)StreamingRecognize supported. Diarization only in BatchRecognize. Speech adaptation up to 1,000 phrases. Language-agnostic auto-detect GA.
chirp_2Older generationNot re-checked in this pass.
telephony / long / short / latest_*Legacy standard modelsPriced as 'Standard' in the V2 table.

Docs and sources

Docs

Sources used

Not fully verified

Whether chirp_3 returns interim results in streaming, chirp_2 status, and any V2 free minutes. The V2 pricing table was read from page markup.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?