GA Microsoft

Azure AI Speech neural and HD voices

Huge catalog (500+ prebuilt voices) with a WebSocket v2 endpoint that accepts streamed text from an LLM via the Speech SDK. DragonHD voices add emotion-aware expressiveness; MAI-Voice is a newer premium tier.

Est. per minute$0.013 - 0.02
1 high-severity warning

Overview

Best for: Enterprises on Azure needing many languages/voices, private networking, containers or disconnected deployment.

At a glance

$/1M chars$15
Free tierYes
Voices500
CloningYes
Text stream inYes
TimestampsYes
EmotionYes
SSMLYes
Latency ms300
WebRTCNo
WebSocketYes
gRPCNo
HIPAAYes
SOC 2Yes
EU dataYes
Self-hostYes
Open weightsNo

Microsoft lists neural and HD voices under 300 ms. $15/1M Neural, $22/1M Neural HD (eastus). Voices: 500+ (DragonHDOmni). Text streaming only via the Speech SDK (C#, C++, Python) on the v2 endpoint, without SSML. Free tier 0.5M chars/month. Cloning via custom neural and gated personal voice. Compliance per Azure scope, not re-verified.

Audio in

Text or SSML (text streaming mode does not support SSML)

Audio out

opus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcm

Languages

Many locales; see language-support page (count not re-verified)

Voices

More than 500 prebuilt; Personal Voice and Professional custom voice (gated access)

Latency

Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms.

Regions

Dozens of Azure regions for standard neural voices; HD voices in a subset

Compliance

Azure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here.

Features

  • input text streaming (Speech SDK, C#/C++/Python)
  • word boundary events
  • visemes
  • custom neural voice
  • personal voice cloning (gated)
  • containers / embedded / disconnected deployment for non-HD voices
  • commitment tiers

Pricing

WhatPriceUnitNotes
Neural (real-time and batch)$15per 1M charactersAzure retail price API, eastus
Neural HD$22per 1M charactersAzure retail price API
Custom neural (professional) real-time$24per 1M charactersPlus $4.032/hour endpoint hosting and training at $52/compute hour
Custom neural HD synthesis$48per 1M characters
Personal Voice synthesis$24per 1M charactersGated feature
Commitment tiers (Neural)$960/80M to $24,000/4,000M per monthmonthlyOverage $12 down to $6 per 1M
How the per-minute estimate was worked out

900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go

Free tier: F0: 0.5M neural characters per month

Source: prices.azure.com

Setup

  1. Create a Speech resource in the Azure portal; copy key and region.
  2. pip install azure-cognitiveservices-speech.
  3. Point SpeechConfig at wss://{region}.tts.speech.microsoft.com/cognitiveservices/websocket/v2.
  4. Create a SpeechSynthesisRequest with input_type TextStream and write LLM chunks into input_stream, then close it.

Endpoint

wss://{region}.tts.speech.microsoft.com/cognitiveservices/websocket/v2

Authentication

Speech resource key (subscription) or Entra ID token

Quick start python

# pip install azure-cognitiveservices-speech
import os
import azure.cognitiveservices.speech as speechsdk

endpoint = f"wss://{os.environ['AZURE_TTS_REGION']}.tts.speech.microsoft.com/cognitiveservices/websocket/v2"
cfg = speechsdk.SpeechConfig(endpoint=endpoint, subscription=os.environ["AZURE_TTS_API_KEY"])
cfg.speech_synthesis_voice_name = "en-US-AvaMultilingualNeural"

synth = speechsdk.SpeechSynthesizer(speech_config=cfg)  # default speaker output
req = speechsdk.SpeechSynthesisRequest(
    input_type=speechsdk.SpeechSynthesisRequestInputType.TextStream)
task = synth.speak_async(req)

for chunk in ["Hello there. ", "This text arrives ", "from an LLM stream."]:
    req.input_stream.write(chunk)
req.input_stream.close()

result = task.get()
print(result.reason)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Text streaming needs the SDK and the v2 endpoint

Text-in streaming only works through the Speech SDK (C#, C++, Python) against /cognitiveservices/websocket/v2. Plain REST or the v1 socket will not accept partial text. No JavaScript support listed.

No SSML in text streaming

When streaming text you set voice and format as global properties; SSML (prosody, breaks, styles) is not supported in that mode.

Azure OpenAI voices excluded

OpenAI voices inside Azure Speech are not supported by text streaming and Microsoft lists them at >500 ms latency.

Custom voice hosting is billed hourly

Professional custom voices cost $4.032 per model per hour to host on top of per-character synthesis; an idle endpoint still costs about $2,900/month.

Pricing page renders without numbers

The public pricing page loads prices dynamically; use the Azure retail price API or calculator to confirm your region.

Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.

Limits

  • Text streaming: SDK only (C#, C++, Python), WebSocket v2 endpoint required
  • HD voices: real-time only, subset of SSML, cloud only
  • Concurrency per resource defaults; see quotas page

Models and products

NameStatusNotes
Prebuilt neural voices (e.g. en-US-AvaMultilingualNeural)GAWork with text streaming on the v2 endpoint.
DragonHD (e.g. en-US-Ava:DragonHDLatestNeural)GA / some Preview voices30+ fine-tuned HD voices, real-time only, subset of SSML.
DragonHDOmniper docs500+ voices with style support.
Azure OpenAI voices in SpeechGANot supported by text streaming; >500 ms latency per Microsoft comparison.
MAI-VoiceListed on pricing pagePrice not resolvable from the public page or retail price API at time of research.

Docs and sources

Docs

Sources used

Not fully verified

MAI-Voice and 'Neural HD Flash' prices (shown as categories on the pricing page but not resolvable); exact locale count.

Similar text-to-speech APIs

Spotted a wrong price or a dead link?