Compare realtime APIs

Pick up to four APIs from any category and see them side by side.

Amazon Transcribe Streaming (incl. Medical and Call Analytics)
AWS
Azure AI Speech real-time speech to text
Microsoft
ElevenLabs Scribe v2 Realtime
ElevenLabs
CategorySpeech-to-textSpeech-to-textSpeech-to-text
StatusGAGAGA
Est. per minute$0.01 - 0.075$0.0067 - 0.025$0.0065 - 0.0098
How that was worked out$0.01/min standard streaming up to $0.075/min Medical. Add PII redaction or CLM as needed.$0.40/hr commitment-tier overage up to $1.20/hr custom + $0.30/hr add-on. Pay-as-you-go standard is $0.0167/min.$0.39/hr base to $0.59/hr with keyterms + editing.
Pricing modelper-minuteper-hourper-hour
Free tier60 minutes per month for 12 months from first request (excludes PII redaction).F0: 5 audio hours per month shared between standard and custom real-time; 1 concurrent request.Free plan includes about 2.5 realtime hours.
Connects byHTTP/2, WebSocketWebSocket (via Speech SDK), REST for short audioWebSocket
Audio inPCM signed 16-bit little-endian (not WAV), FLAC, or Opus in Ogg. 16 kHz recommended; 8 kHz telephony accepted. Chunks of 50-200 ms recommended; send zero-byte silence rather than pausing.Speech SDK handles microphone, files and push/pull streams; default 16 kHz 16-bit mono PCM; compressed formats via GStreamer.audio_format: pcm_8000, pcm_16000 (default), pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000; audio sent base64 in JSON input_audio_chunk messages.
Audio outn/an/an/a
LanguagesMany streaming languages (count not re-verified); Medical is US English.100+ locales (see language support page; count not re-verified).90+ languages (vendor); language_code plus secondary_languages hints.
Latency (vendor claim)No vendor figure captured; latency depends on chunk size per docs.No vendor figure captured.Vendor claim: ~150 ms (footnoted, conditions not stated).
Key limits
  • 25 concurrent standard streams per region by default (adjustable); Medical and Call Analytics streams also 25.
  • Streams must be close to real time; LimitExceededException on overuse.
  • Billed in 1-second increments, no per-request minimum for transcription.
  • Concurrent real-time requests: F0 1 (fixed), S0 100 default (adjustable); limit is shared with speech translation.
  • Diarization identifies up to 35 speakers (errors beyond that).
  • Realtime concurrency limit by plan is in a separate chart on the models page (not captured).
  • keepalive_interval_ms configurable 500-10000; errors include session_time_limit_exceeded, commit_throttled, queue_overflow, insufficient_audio_activity.
High-severity warnings
  • Old price figures circulate
  • Pricing page shows no numbers
  • None
ComplianceAmazon Transcribe and Transcribe Medical are HIPAA-eligible under the AWS BAA per AWS (not re-verified here). AI services opt-out policy controls data use for training.Azure compliance scope (HIPAA BAA, SOC, ISO) applies to Azure AI Speech per Microsoft; connected/disconnected containers available. Not re-verified.Data-residency hosts for EU, India, Singapore; enable_logging=false for zero retention. Certifications not re-verified.
Self-hostableNoYesNo
Last checked2026-10-102026-10-102026-10-10
Key numbers and features
$/hour$0.6$1$0.39
Bills silenceYes--
Free tierYesYesYes
Free credit $---
Live speakersYesYes-
Turn detect--Yes
KeytermsYesYesYes
PII redactYes--
PartialsYesYesYes
Mixed langs-Yes-
8 kHz phoneNo-Yes
Latency ms--150
Languages77-90
Max session min---
Concurrency251009
WebRTCNoNoNo
WebSocketYesYesYes
gRPCNoNoNo
HIPAAYesYes-
SOC 2-Yes-
EU dataYesYesYes
Self-hostNoYesNo
Open weightsNoNoNo
High warnings110