Amazon Transcribe Streaming (incl. Medical and Call Analytics)
Streaming STT over HTTP/2 or WebSocket with SigV4 auth, plus Transcribe Medical streaming and real-time Call Analytics. Current us-east-1 price list shows a flat $0.01/min for standard streaming.
Overview
Best for: AWS-centric call centers and healthcare (Medical, Call Analytics) needing IAM, VPC and HIPAA-eligible services.
At a glance
$0.01/min us-east-1 price list (older articles quote $0.024/min). 77 languages support streaming, some not in every region. Free tier 60 min/month for 12 months. Accepts 8 kHz PCM, FLAC or Opus but not mulaw. AWS asks you to send silence rather than gaps, so silence is billed. Also HTTP/2 transport. EU via AWS EU regions.
PCM signed 16-bit little-endian (not WAV), FLAC, or Opus in Ogg. 16 kHz recommended; 8 kHz telephony accepted. Chunks of 50-200 ms recommended; send zero-byte silence rather than pausing.
n/a
Many streaming languages (count not re-verified); Medical is US English.
No vendor figure captured; latency depends on chunk size per docs.
Most commercial AWS regions plus GovCloud; FIPS endpoints in US regions.
Amazon Transcribe and Transcribe Medical are HIPAA-eligible under the AWS BAA per AWS (not re-verified here). AI services opt-out policy controls data use for training.
Features
- partial results with stabilization
- custom vocabulary and vocabulary filters
- custom language models (CLM, priced separately)
- PII identification/redaction (paid)
- speaker partitioning
- channel identification
- automatic language identification
- toxicity detection (batch)
- Call Analytics real-time insights
Pricing
| What | Price | Unit |
|---|---|---|
| Standard streaming (us-east-1) | $0.01 | per minute ($0.0001667/s) |
| PII redaction streaming | +$0.0024 / $0.0015 / $0.00102 / $0.00078 | per minute by tier (0-250K, 250K-1M, 1M-5M, 5M+) |
| Custom language model streaming | $0.006 first 250K | per minute ($0.0001/s), tiered down |
| Medical streaming | $0.075 | per minute ($0.00125/s) |
| Call Analytics streaming | $0.030 / $0.0186 / $0.0138 / $0.0114 | per minute by tier |
| Batch (reference) | $0.006 | per minute |
$0.01/min standard streaming up to $0.075/min Medical. Add PII redaction or CLM as needed.
Free tier: 60 minutes per month for 12 months from first request (excludes PII redaction).
Source: aws.amazon.com
Setup
- Create IAM credentials with transcribe:StartStreamTranscription.
- Use an AWS SDK with event-stream support (JS v3 @aws-sdk/client-transcribe-streaming, Java, Go, or the Python amazon-transcribe package).
- Start a stream with language, sample rate and encoding, push audio events, read TranscriptEvents (IsPartial flag).
- Browsers: use WebSocket with a SigV4 presigned URL generated server-side, or Cognito credentials.
Endpoint
transcribestreaming.<region>.amazonaws.com (HTTP/2 :443; WebSocket :8443 /stream-transcription-websocket with presigned URL)
Authentication
AWS SigV4 (IAM). WebSocket uses a presigned URL (max 5 minute validity).
Quick start python
import asyncio # pip install amazon-transcribe
from amazon_transcribe.client import TranscribeStreamingClient
from amazon_transcribe.handlers import TranscriptResultStreamHandler
class Printer(TranscriptResultStreamHandler):
async def handle_transcript_event(self, event):
for r in event.transcript.results:
for alt in r.alternatives:
print("partial" if r.is_partial else "FINAL", alt.transcript)
async def main():
client = TranscribeStreamingClient(region="us-east-1")
stream = await client.start_stream_transcription(
language_code="en-US", media_sample_rate_hz=16000, media_encoding="pcm")
async def write():
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(3200): # 100 ms
await stream.input_stream.send_audio_event(audio_chunk=chunk)
await asyncio.sleep(0.1)
await stream.input_stream.end_stream()
await asyncio.gather(write(), Printer(stream.output_stream).handle_events())
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Old price figures circulate
Many articles quote $0.024/min tiered streaming. The current AWS price list (us-east-1) shows $0.0001667/s = $0.01/min for StreamingAudio. Prices vary by region; check the price list API for yours.
Send silence, not gaps
AWS says to send zero-byte PCM silence when there is no speech and to keep 50-200 ms chunks. Gaps or bursty sends raise latency or trigger errors.
Medical is 7.5x standard
Transcribe Medical streaming is $0.075/min and US English only; Call Analytics streaming starts at $0.03/min.
WAV headers are not PCM
Streaming PCM must be raw signed 16-bit little-endian. Sending a WAV file including its header corrupts the start of the transcript.
Python SDK is a separate package
boto3 does not do streaming transcription; use the amazon-transcribe package or another SDK with HTTP/2 event streams.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- 25 concurrent standard streams per region by default (adjustable); Medical and Call Analytics streams also 25.
- Streams must be close to real time; LimitExceededException on overuse.
- Billed in 1-second increments, no per-request minimum for transcription.
Models and products
| Name | Status |
|---|---|
| Standard streaming | GA |
| Transcribe Medical streaming | GA |
| Call Analytics streaming | GA |
Docs and sources
Docs
Sources used
- aws.amazon.com/transcribe/pricing/
- pricing.us-east-1.amazonaws.com/offers/v1.0/aws/transcribe/current/us-east-1/in...
- docs.aws.amazon.com/transcribe/latest/dg/streaming.html
- docs.aws.amazon.com/general/latest/gr/transcribe.html
The flat $0.01/min streaming price is from the price list JSON and the page's worked example; the rendered tier table was not visible. PII tier 4 value inferred from the price list (5M+ = $0.000013/s). Streaming language count.