Azure AI Speech neural and HD voices
Huge catalog (500+ prebuilt voices) with a WebSocket v2 endpoint that accepts streamed text from an LLM via the Speech SDK. DragonHD voices add emotion-aware expressiveness; MAI-Voice is a newer premium tier.
Overview
Best for: Enterprises on Azure needing many languages/voices, private networking, containers or disconnected deployment.
At a glance
Microsoft lists neural and HD voices under 300 ms. $15/1M Neural, $22/1M Neural HD (eastus). Voices: 500+ (DragonHDOmni). Text streaming only via the Speech SDK (C#, C++, Python) on the v2 endpoint, without SSML. Free tier 0.5M chars/month. Cloning via custom neural and gated personal voice. Compliance per Azure scope, not re-verified.
Text or SSML (text streaming mode does not support SSML)
opus, mp3, pcm, truesilk at 8/16/24/48 kHz; raw PCM formats such as Raw24Khz16BitMonoPcm
Many locales; see language-support page (count not re-verified)
More than 500 prebuilt; Personal Voice and Professional custom voice (gated access)
Microsoft comparison table: HD and standard neural voices < 300 ms; Azure OpenAI voices > 500 ms.
Dozens of Azure regions for standard neural voices; HD voices in a subset
Azure compliance programs; containers and disconnected options for non-HD voices. Specific certifications not re-verified here.
Features
- input text streaming (Speech SDK, C#/C++/Python)
- word boundary events
- visemes
- custom neural voice
- personal voice cloning (gated)
- containers / embedded / disconnected deployment for non-HD voices
- commitment tiers
Pricing
| What | Price | Unit |
|---|---|---|
| Neural (real-time and batch) | $15 | per 1M characters |
| Neural HD | $22 | per 1M characters |
| Custom neural (professional) real-time | $24 | per 1M characters |
| Custom neural HD synthesis | $48 | per 1M characters |
| Personal Voice synthesis | $24 | per 1M characters |
| Commitment tiers (Neural) | $960/80M to $24,000/4,000M per month | monthly |
900 chars/min; Neural $15/1M vs Neural HD $22/1M at pay-as-you-go
Free tier: F0: 0.5M neural characters per month
Source: prices.azure.com
Setup
- Create a Speech resource in the Azure portal; copy key and region.
- pip install azure-cognitiveservices-speech.
- Point SpeechConfig at wss://{region}.tts.speech.microsoft.com/cognitiveservices/websocket/v2.
- Create a SpeechSynthesisRequest with input_type TextStream and write LLM chunks into input_stream, then close it.
Endpoint
wss://{region}.tts.speech.microsoft.com/cognitiveservices/websocket/v2
Authentication
Speech resource key (subscription) or Entra ID token
Quick start python
# pip install azure-cognitiveservices-speech
import os
import azure.cognitiveservices.speech as speechsdk
endpoint = f"wss://{os.environ['AZURE_TTS_REGION']}.tts.speech.microsoft.com/cognitiveservices/websocket/v2"
cfg = speechsdk.SpeechConfig(endpoint=endpoint, subscription=os.environ["AZURE_TTS_API_KEY"])
cfg.speech_synthesis_voice_name = "en-US-AvaMultilingualNeural"
synth = speechsdk.SpeechSynthesizer(speech_config=cfg) # default speaker output
req = speechsdk.SpeechSynthesisRequest(
input_type=speechsdk.SpeechSynthesisRequestInputType.TextStream)
task = synth.speak_async(req)
for chunk in ["Hello there. ", "This text arrives ", "from an LLM stream."]:
req.input_stream.write(chunk)
req.input_stream.close()
result = task.get()
print(result.reason)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Text streaming needs the SDK and the v2 endpoint
Text-in streaming only works through the Speech SDK (C#, C++, Python) against /cognitiveservices/websocket/v2. Plain REST or the v1 socket will not accept partial text. No JavaScript support listed.
No SSML in text streaming
When streaming text you set voice and format as global properties; SSML (prosody, breaks, styles) is not supported in that mode.
Azure OpenAI voices excluded
OpenAI voices inside Azure Speech are not supported by text streaming and Microsoft lists them at >500 ms latency.
Custom voice hosting is billed hourly
Professional custom voices cost $4.032 per model per hour to host on top of per-character synthesis; an idle endpoint still costs about $2,900/month.
Pricing page renders without numbers
The public pricing page loads prices dynamically; use the Azure retail price API or calculator to confirm your region.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Text streaming: SDK only (C#, C++, Python), WebSocket v2 endpoint required
- HD voices: real-time only, subset of SSML, cloud only
- Concurrency per resource defaults; see quotas page
Models and products
| Name | Status |
|---|---|
| Prebuilt neural voices (e.g. en-US-AvaMultilingualNeural) | GA |
| DragonHD (e.g. en-US-Ava:DragonHDLatestNeural) | GA / some Preview voices |
| DragonHDOmni | per docs |
| Azure OpenAI voices in Speech | GA |
| MAI-Voice | Listed on pricing page |
Docs and sources
Docs
Sources used
- prices.azure.com/api/retail/prices
- azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/
- learn.microsoft.com/en-us/azure/ai-services/speech-service/how-to-lower-speech-...
- learn.microsoft.com/en-us/azure/ai-services/speech-service/high-definition-voic...
MAI-Voice and 'Neural HD Flash' prices (shown as categories on the pricing page but not resolvable); exact locale count.