GA Amazon Web Services

Amazon Polly (generative and bidirectional streaming)

Polly now has StartSpeechSynthesisStream, an HTTP/2 bidirectional API that takes text incrementally and returns audio as it is produced, but only for the generative engine. Standard and neural engines remain request/response with streamed output.

Est. per minute$0.014 - 0.027
1 high-severity warning

Overview

Best for: AWS-native stacks (Connect, Lex) wanting IAM auth and predictable per-character pricing.

At a glance

$/1M chars$30
Free tierYes
Free credit $$200
CloningNo
Instant cloneNo
Text stream inYes
TimestampsYes
SSMLYes
8 kHz phoneNo
WebRTCNo
WebSocketNo
gRPCNo
EU dataYes
Self-hostNo
Open weightsNo

$30/1M is the generative engine, the only one that accepts bidirectional text streaming (neural $16/1M). Transport is an HTTP/2 event stream. Stream API covers about 40 locales. Speech marks (word timings) only on the non-stream API. Free tier is 100K generative chars/month for 12 months; up to $200 AWS credits for new accounts. EU region per third-party list.

Audio in

Plain text or SSML

Audio out

mp3, ogg_opus, ogg_vorbis, pcm (stream API; JSON speech marks not supported on the stream API)

Languages

Stream API LanguageCode list covers ~40 locales (only needed for bilingual voices)

Voices

Generative voices subset of Polly catalog; no voice cloning

Latency

No numeric claim found.

Regions

Bidi streaming region list not confirmed officially; a third-party package lists us-east-1, us-west-2, eu-central-1, eu-west-2, ap-southeast-1, ca-central-1 (2026-05)

Compliance

AWS compliance programs (GovCloud availability for standard/neural).

Features

  • bidirectional text-in streaming (generative only)
  • flush via FlushStreamConfiguration
  • lexicons (up to 5)
  • speech marks (non-stream API)

Pricing

WhatPriceUnitNotes
Generative voices$30per 1M characters
Neural voices$16per 1M characters
Standard voices$4per 1M characters
Long-form voices$100per 1M characters
How the per-minute estimate was worked out

900 chars/min. Bidi streaming requires generative ($30/1M = $0.027); neural $16/1M = $0.0144 without input streaming.

Free tier: 12 months: 5M standard, 1M neural, 500K long-form, 100K generative characters per month; new accounts also get up to $200 AWS credits

Source: aws.amazon.com

Setup

  1. Create IAM credentials with polly:SynthesizeSpeech (and the stream action).
  2. For true text-in streaming use an SDK that exposes StartSpeechSynthesisStream (AWS Java SDK per third-party notes; boto3 support unconfirmed).
  3. Otherwise call SynthesizeSpeech per sentence with Engine=generative and read the AudioStream.

Endpoint

POST /v1/synthesisStream (StartSpeechSynthesisStream); SynthesizeSpeech for request/response

Authentication

AWS SigV4 (IAM credentials)

Quick start python

# pip install boto3  - per-sentence request with streamed output
import boto3

polly = boto3.client("polly", region_name="us-east-1")
resp = polly.synthesize_speech(
    Text="Hello from Amazon Polly's generative engine.",
    VoiceId="Ruth",            # pick a voice that supports the generative engine
    Engine="generative",
    OutputFormat="pcm",
    SampleRate="16000",
)
with open("out_16k_s16le.pcm", "wb") as f:
    for chunk in resp["AudioStream"].iter_chunks():
        f.write(chunk)
# For text-in streaming use StartSpeechSynthesisStream (generative only)
# from an SDK that supports HTTP/2 event streams.

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Bidi streaming is generative-only

StartSpeechSynthesisStream accepts only the generative engine even though the parameter lists others. Generative costs $30/1M chars, nearly 2x neural.

SDK support is uneven

A third-party package says boto3 does not expose StartSpeechSynthesisStream and only the Java SDK does. Check your SDK version before designing around it.

Generative free tier is tiny

Only 100K generative characters/month for 12 months, about 2 hours of audio.

No voice cloning

Polly has no self-serve cloning; brand voices are an enterprise engagement.

Speech marks not on the stream API

JSON speech marks (word timings) are not supported by the bidi stream; use separate SynthesizeSpeech calls if you need them.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Bidi stream: generative engine only
  • Up to 5 lexicons
  • GovCloud prices differ ($4.80 standard, $19.20 neural)

Models and products

NameStatusNotes
generative engineGAOnly engine accepted by StartSpeechSynthesisStream.
long-form engineGA$100/1M chars; not for real-time.
neural engineGASynthesizeSpeech with streamed output.
standard engineGACheapest, oldest.

Docs and sources

Docs

Sources used

Not fully verified

Which SDKs support the bidi stream; region list; whether 'Ruth' is the right generative voice for your locale.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?