GA Microsoft

Azure AI Speech real-time speech to text

Mature real-time STT via the Speech SDK (WebSocket under the hood) with intermediate results, phrase lists, custom speech models, continuous language ID and real-time diarization add-ons.

Est. per minute$0.0067 - 0.025
1 high-severity warning

Overview

Best for: Enterprises on Azure needing custom acoustic/language models, on-prem containers, and Microsoft compliance coverage.

At a glance

$/hour$1
Free tierYes
Live speakersYes
KeytermsYes
PartialsYes
Mixed langsYes
Concurrency100
WebRTCNo
WebSocketYes
gRPCNo
HIPAAYes
SOC 2Yes
EU dataYes
Self-hostYes
Open weightsNo

100+ locales, count not verified. $1.00/hr is eastus PAYG from the retail price API; commitment tiers go down to $0.40/hr. Diarization and continuous language ID are +$0.30/hr add-ons each. Free tier F0 is 5 audio hours/month with 1 concurrent stream. Compliance per Azure scope, not re-verified. WebSocket via Speech SDK.

Audio in

Speech SDK handles microphone, files and push/pull streams; default 16 kHz 16-bit mono PCM; compressed formats via GStreamer.

Audio out

n/a

Languages

100+ locales (see language support page; count not re-verified).

Latency

No vendor figure captured.

Regions

Many Azure regions; prices above are eastus retail list prices.

Compliance

Azure compliance scope (HIPAA BAA, SOC, ISO) applies to Azure AI Speech per Microsoft; connected/disconnected containers available. Not re-verified.

Features

  • intermediate (recognizing) results
  • continuous recognition
  • phrase lists
  • custom speech models
  • continuous language identification (paid add-on)
  • real-time diarization (paid add-on)
  • pronunciation assessment
  • multichannel (preview)

Pricing

WhatPriceUnitNotes
Standard real-time STT (S1, eastus)$1.00per audio hourAzure retail prices API meter 'S1 Speech To Text'.
Custom real-time STT$1.20per audio hourPlus custom endpoint hosting ~$0.0538/hr per model.
Real-time enhanced features add-on (diarization / language ID)$0.30per audio hour per featureMeter 'S1 Speech to Text Enhanced Feature Audio'.
Commitment tiers2K / 10K / 50K / 100K hour bundlesmonthlyOverage $0.80/hr (2K), $0.65 (10K), $0.50 (50K), $0.40 (100K) for standard STT.
Batch STT (reference)$0.18per audio hourFast transcription $0.36/hr.
How the per-minute estimate was worked out

$0.40/hr commitment-tier overage up to $1.20/hr custom + $0.30/hr add-on. Pay-as-you-go standard is $0.0167/min.

Free tier: F0: 5 audio hours per month shared between standard and custom real-time; 1 concurrent request.

Source: azure.microsoft.com

Setup

  1. Create a Speech (or Foundry) resource and copy key + region.
  2. pip install azure-cognitiveservices-speech.
  3. Create SpeechConfig and a SpeechRecognizer; subscribe to recognizing (partial) and recognized (final) events.
  4. Call start_continuous_recognition; for browsers fetch a short-lived authorization token from your backend (issueToken endpoint).

Endpoint

Speech SDK manages the WebSocket (wss://<region>.stt.speech.microsoft.com/...); use SDK rather than raw protocol.

Authentication

Ocp-Apim-Subscription-Key / resource key, Microsoft Entra ID, or 10-minute authorization tokens for client apps.

Quick start python

import os, time
import azure.cognitiveservices.speech as speechsdk  # pip install azure-cognitiveservices-speech

cfg = speechsdk.SpeechConfig(subscription=os.environ["SPEECH_KEY"],
                             region=os.environ["SPEECH_REGION"])
cfg.speech_recognition_language = "en-US"
audio = speechsdk.audio.AudioConfig(use_default_microphone=True)
rec = speechsdk.SpeechRecognizer(speech_config=cfg, audio_config=audio)

rec.recognizing.connect(lambda e: print("partial", e.result.text))
rec.recognized.connect(lambda e: print("FINAL", e.result.text))
rec.canceled.connect(lambda e: print("canceled", e.cancellation_details))

rec.start_continuous_recognition()
time.sleep(30)
rec.stop_continuous_recognition()

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Pricing page shows no numbers

The public Speech pricing page renders rates as placeholders unless a region/currency loads; the figures here come from the Azure retail prices API (eastus, USD). Re-check for your region.

Most expensive mainstream list price

$1.00 per audio hour pay-as-you-go is several times Deepgram, AssemblyAI or Soniox. Commitment tiers are needed to get near $0.40-0.50/hr.

Diarization and language ID cost extra

Real-time diarization and continuous language identification are billed as add-ons (about $0.30 per audio hour per feature).

Free tier is 1 concurrent stream

F0 allows a single concurrent real-time request and the limit cannot be raised; S0 starts at 100.

Custom models carry hosting fees

Custom speech endpoints bill hosting per model per hour even when idle; unused free-tier models are decommissioned after 7 days.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Concurrent real-time requests: F0 1 (fixed), S0 100 default (adjustable); limit is shared with speech translation.
  • Diarization identifies up to 35 speakers (errors beyond that).

Models and products

NameStatusNotes
Standard (base) real-time modelGALatest base model per locale used by default.
Custom speech endpointGATrained/adapted model, needs hosted endpoint (hourly hosting fee).
MAI-Transcribe-2 / MAI-Transcribe-1.5Listed on pricing pagePriced per hour; MAI-Transcribe-2 under a limited-time promotion to 2026-12-31. Real-time availability not confirmed.
Multichannel real-time transcriptionPreviewUp to two channels tagged independently.

Docs and sources

Docs

Sources used

Not fully verified

Mapping of the 'Enhanced Feature Audio' meter to diarization/language-ID add-ons is inferred from meter names; MAI-Transcribe real-time support and prices; locale count.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?