Amazon Polly (generative and bidirectional streaming)
Polly now has StartSpeechSynthesisStream, an HTTP/2 bidirectional API that takes text incrementally and returns audio as it is produced, but only for the generative engine. Standard and neural engines remain request/response with streamed output.
Overview
Best for: AWS-native stacks (Connect, Lex) wanting IAM auth and predictable per-character pricing.
At a glance
$30/1M is the generative engine, the only one that accepts bidirectional text streaming (neural $16/1M). Transport is an HTTP/2 event stream. Stream API covers about 40 locales. Speech marks (word timings) only on the non-stream API. Free tier is 100K generative chars/month for 12 months; up to $200 AWS credits for new accounts. EU region per third-party list.
Plain text or SSML
mp3, ogg_opus, ogg_vorbis, pcm (stream API; JSON speech marks not supported on the stream API)
Stream API LanguageCode list covers ~40 locales (only needed for bilingual voices)
Generative voices subset of Polly catalog; no voice cloning
No numeric claim found.
Bidi streaming region list not confirmed officially; a third-party package lists us-east-1, us-west-2, eu-central-1, eu-west-2, ap-southeast-1, ca-central-1 (2026-05)
AWS compliance programs (GovCloud availability for standard/neural).
Features
- bidirectional text-in streaming (generative only)
- flush via FlushStreamConfiguration
- lexicons (up to 5)
- speech marks (non-stream API)
Pricing
| What | Price | Unit |
|---|---|---|
| Generative voices | $30 | per 1M characters |
| Neural voices | $16 | per 1M characters |
| Standard voices | $4 | per 1M characters |
| Long-form voices | $100 | per 1M characters |
900 chars/min. Bidi streaming requires generative ($30/1M = $0.027); neural $16/1M = $0.0144 without input streaming.
Free tier: 12 months: 5M standard, 1M neural, 500K long-form, 100K generative characters per month; new accounts also get up to $200 AWS credits
Source: aws.amazon.com
Setup
- Create IAM credentials with polly:SynthesizeSpeech (and the stream action).
- For true text-in streaming use an SDK that exposes StartSpeechSynthesisStream (AWS Java SDK per third-party notes; boto3 support unconfirmed).
- Otherwise call SynthesizeSpeech per sentence with Engine=generative and read the AudioStream.
Endpoint
POST /v1/synthesisStream (StartSpeechSynthesisStream); SynthesizeSpeech for request/response
Authentication
AWS SigV4 (IAM credentials)
Quick start python
# pip install boto3 - per-sentence request with streamed output
import boto3
polly = boto3.client("polly", region_name="us-east-1")
resp = polly.synthesize_speech(
Text="Hello from Amazon Polly's generative engine.",
VoiceId="Ruth", # pick a voice that supports the generative engine
Engine="generative",
OutputFormat="pcm",
SampleRate="16000",
)
with open("out_16k_s16le.pcm", "wb") as f:
for chunk in resp["AudioStream"].iter_chunks():
f.write(chunk)
# For text-in streaming use StartSpeechSynthesisStream (generative only)
# from an SDK that supports HTTP/2 event streams.
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Bidi streaming is generative-only
StartSpeechSynthesisStream accepts only the generative engine even though the parameter lists others. Generative costs $30/1M chars, nearly 2x neural.
SDK support is uneven
A third-party package says boto3 does not expose StartSpeechSynthesisStream and only the Java SDK does. Check your SDK version before designing around it.
Generative free tier is tiny
Only 100K generative characters/month for 12 months, about 2 hours of audio.
No voice cloning
Polly has no self-serve cloning; brand voices are an enterprise engagement.
Speech marks not on the stream API
JSON speech marks (word timings) are not supported by the bidi stream; use separate SynthesizeSpeech calls if you need them.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Bidi stream: generative engine only
- Up to 5 lexicons
- GovCloud prices differ ($4.80 standard, $19.20 neural)
Models and products
| Name | Status |
|---|---|
| generative engine | GA |
| long-form engine | GA |
| neural engine | GA |
| standard engine | GA |
Docs and sources
Docs
Sources used
- aws.amazon.com/polly/pricing/
- docs.aws.amazon.com/polly/latest/dg/API_StartSpeechSynthesisStream.html
- amazon-polly-streaming.readthedocs.io/en/stable/overview.html
Which SDKs support the bidi stream; region list; whether 'Ruth' is the right generative voice for your locale.