Azure AI Speech real-time speech to text
Mature real-time STT via the Speech SDK (WebSocket under the hood) with intermediate results, phrase lists, custom speech models, continuous language ID and real-time diarization add-ons.
Overview
Best for: Enterprises on Azure needing custom acoustic/language models, on-prem containers, and Microsoft compliance coverage.
At a glance
100+ locales, count not verified. $1.00/hr is eastus PAYG from the retail price API; commitment tiers go down to $0.40/hr. Diarization and continuous language ID are +$0.30/hr add-ons each. Free tier F0 is 5 audio hours/month with 1 concurrent stream. Compliance per Azure scope, not re-verified. WebSocket via Speech SDK.
Speech SDK handles microphone, files and push/pull streams; default 16 kHz 16-bit mono PCM; compressed formats via GStreamer.
n/a
100+ locales (see language support page; count not re-verified).
No vendor figure captured.
Many Azure regions; prices above are eastus retail list prices.
Azure compliance scope (HIPAA BAA, SOC, ISO) applies to Azure AI Speech per Microsoft; connected/disconnected containers available. Not re-verified.
Features
- intermediate (recognizing) results
- continuous recognition
- phrase lists
- custom speech models
- continuous language identification (paid add-on)
- real-time diarization (paid add-on)
- pronunciation assessment
- multichannel (preview)
Pricing
| What | Price | Unit |
|---|---|---|
| Standard real-time STT (S1, eastus) | $1.00 | per audio hour |
| Custom real-time STT | $1.20 | per audio hour |
| Real-time enhanced features add-on (diarization / language ID) | $0.30 | per audio hour per feature |
| Commitment tiers | 2K / 10K / 50K / 100K hour bundles | monthly |
| Batch STT (reference) | $0.18 | per audio hour |
$0.40/hr commitment-tier overage up to $1.20/hr custom + $0.30/hr add-on. Pay-as-you-go standard is $0.0167/min.
Free tier: F0: 5 audio hours per month shared between standard and custom real-time; 1 concurrent request.
Source: azure.microsoft.com
Setup
- Create a Speech (or Foundry) resource and copy key + region.
- pip install azure-cognitiveservices-speech.
- Create SpeechConfig and a SpeechRecognizer; subscribe to recognizing (partial) and recognized (final) events.
- Call start_continuous_recognition; for browsers fetch a short-lived authorization token from your backend (issueToken endpoint).
Endpoint
Speech SDK manages the WebSocket (wss://<region>.stt.speech.microsoft.com/...); use SDK rather than raw protocol.
Authentication
Ocp-Apim-Subscription-Key / resource key, Microsoft Entra ID, or 10-minute authorization tokens for client apps.
Quick start python
import os, time
import azure.cognitiveservices.speech as speechsdk # pip install azure-cognitiveservices-speech
cfg = speechsdk.SpeechConfig(subscription=os.environ["SPEECH_KEY"],
region=os.environ["SPEECH_REGION"])
cfg.speech_recognition_language = "en-US"
audio = speechsdk.audio.AudioConfig(use_default_microphone=True)
rec = speechsdk.SpeechRecognizer(speech_config=cfg, audio_config=audio)
rec.recognizing.connect(lambda e: print("partial", e.result.text))
rec.recognized.connect(lambda e: print("FINAL", e.result.text))
rec.canceled.connect(lambda e: print("canceled", e.cancellation_details))
rec.start_continuous_recognition()
time.sleep(30)
rec.stop_continuous_recognition()
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Pricing page shows no numbers
The public Speech pricing page renders rates as placeholders unless a region/currency loads; the figures here come from the Azure retail prices API (eastus, USD). Re-check for your region.
Most expensive mainstream list price
$1.00 per audio hour pay-as-you-go is several times Deepgram, AssemblyAI or Soniox. Commitment tiers are needed to get near $0.40-0.50/hr.
Diarization and language ID cost extra
Real-time diarization and continuous language identification are billed as add-ons (about $0.30 per audio hour per feature).
Free tier is 1 concurrent stream
F0 allows a single concurrent real-time request and the limit cannot be raised; S0 starts at 100.
Custom models carry hosting fees
Custom speech endpoints bill hosting per model per hour even when idle; unused free-tier models are decommissioned after 7 days.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Concurrent real-time requests: F0 1 (fixed), S0 100 default (adjustable); limit is shared with speech translation.
- Diarization identifies up to 35 speakers (errors beyond that).
Models and products
| Name | Status |
|---|---|
| Standard (base) real-time model | GA |
| Custom speech endpoint | GA |
| MAI-Transcribe-2 / MAI-Transcribe-1.5 | Listed on pricing page |
| Multichannel real-time transcription | Preview |
Docs and sources
Docs
Sources used
- azure.microsoft.com/en-us/pricing/details/speech/
- prices.azure.com/api/retail/prices?$filter=armRegionName eq 'eastus' and...
- learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-to-text
- learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-services-quot...
Mapping of the 'Enhanced Feature Audio' meter to diarization/language-ID add-ons is inferred from meter names; MAI-Transcribe real-time support and prices; locale count.