Azure Speech Voice Live API
Managed voice-agent API in Microsoft Foundry that wraps either native speech-to-speech models (gpt-realtime family, azure-realtime, phi4-mm-realtime) or text LLMs with Azure speech-to-text and 600+ Azure TTS voices, adding noise suppression, server echo cancellation, semantic VAD and avatars. Suits contact centers and branded agents on Azure.
Overview
Best for: Contact centers and branded assistants on Azure that want noise suppression, server echo cancellation, multilingual semantic VAD, hundreds of voices, custom voice or a talking avatar without stitching services together.
At a glance
Supports both native speech-to-speech models and a cascaded STT+LLM+TTS mode. 146 input / 151 output locales. 600+ Azure neural voices; custom voice is limited access. Prices are Pro tier with gpt-realtime-2.1 (Lite from $4/$12). HIPAA via Microsoft BAA for in-scope services; confirm Voice Live scope. Telephony via Azure Communication Services, not direct SIP. BYO model is preview.
PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session.
PCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264).
146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi.
600+ Azure neural voices across 150+ locales, 30+ HD (DragonHD) voices, MAI-Voice-2-Flash (preview), 34 azure-realtime-native voices, native gpt-realtime voices, and custom voice (limited access).
No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models.
10+ Azure regions; LLM availability and processing scope (global, data zone, regional) depend on the resource region and the model suffix.
Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft).
Features
- azure_semantic_vad and azure_semantic_vad_multilingual turn detection (with filler-word removal), plus server_vad and semantic_vad
- azure_deep_noise_suppression
- server_echo_cancellation, plus Live-Reference AEC (client-supplied playback reference, API 2026-07-15+)
- function calling incl. asynchronous calls; MCP in model mode (not with phi models)
- Foundry Agent Service integration (agent_id + project_id)
- phrase lists and custom speech models for recognition
- custom voice, HD voice temperature, speaking rate, custom lexicon
- word-level audio timestamps and visemes
- text-to-speech avatars, photo avatars (VASA-1, VASA-2 preview), custom avatars
- transcription models: azure-speech, mai-transcribe-2 (preview), whisper-1, gpt-4o-transcribe(-diarize)
- telephony through Azure Communication Services
Pricing
| What | Price | Unit |
|---|---|---|
| Pro: LLM audio input / output (e.g. gpt-realtime-2.1 native audio) | $32.00 / $64.00 | per 1M tokens |
| Pro: LLM text input / output | $4.00 / $16.00 (a second Pro text-output meter is $24.00) | per 1M tokens |
| Pro: Azure standard speech audio input / output | $17.00 / $31.00 | per 1M tokens |
| Pro: Azure custom speech audio input / output | $40.00 / $55.00 | per 1M tokens |
| Standard: LLM audio input / output (e.g. gpt-realtime-2.1-mini) | $11.00 / $22.00 | per 1M tokens |
| Standard: LLM text input / output | $0.66 / $2.64 | per 1M tokens |
| Standard: Azure standard speech audio input / output | $15.00 / $26.00 | per 1M tokens |
| Standard: Azure custom speech audio input / output | $39.00 / $50.00 | per 1M tokens |
| Lite: LLM audio input / output (phi4-mm-realtime) | $4.00 / $12.00 | per 1M tokens |
| Lite: LLM text input / output | $0.11 / $0.44 | per 1M tokens |
| Lite: Azure standard speech audio input / output | $15.00 / $25.00 | per 1M tokens |
| Lite: Azure custom speech audio input / output | $38.00 / $50.00 | per 1M tokens |
| BYO model: standard speech audio input / output | $12.50 / $23.00 | per 1M tokens |
| Avatars, custom voice training and hosting | separate | Speech service pricing |
Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra.
Microsoft: Azure OpenAI models ~10 tokens/s input audio and ~20 tokens/s output audio; Phi models ~12.5 tokens/s input, ~20 tokens/s output. Token rate for Azure speech (STT/TTS) meters is not published.
Free tier: No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request).
Source: learn.microsoft.com
Setup
- Create an Azure subscription and a Microsoft Foundry resource (recommended over a plain Speech resource; Speech resources lack Agent Service and BYOM).
- Assign Cognitive Services User and Foundry User roles to your identity, or copy the resource key.
- Pick a model by name (it sets the Pro / Standard / Lite price); no deployment is needed for natively supported models.
- Connect a server WebSocket to the voice-live/realtime endpoint with api-version and model query parameters.
- Send session.update with turn detection, noise suppression, echo cancellation and voice, then stream input_audio_buffer.append.
- For browsers use the Voice Live WebRTC flow or relay through your backend; for avatars complete the session.avatar.connect SDP exchange.
Endpoint
wss://<foundry-resource>.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10&model=gpt-realtime-2.1 (older resources: <name>.cognitiveservices.azure.com)
Authentication
Recommended Microsoft Entra token (scope https://ai.azure.com/.default) as Authorization: Bearer. API key works as an api-key header (server only) or an api-key query parameter; never use the query-string key in a browser.
Quick start javascript
import WebSocket from "ws"; // server side (API key in header is not possible from a browser)
const RES = process.env.FOUNDRY_RESOURCE; // Foundry resource name
const url = `wss://${RES}.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10&model=gpt-realtime-2.1`;
const ws = new WebSocket(url, { headers: { "api-key": process.env.FOUNDRY_API_KEY } }); // or Authorization: Bearer <Entra token>
ws.on("open", () => {
ws.send(JSON.stringify({
type: "session.update",
session: {
instructions: "You are a friendly support agent. Keep answers short.",
input_audio_sampling_rate: 24000,
turn_detection: { type: "azure_semantic_vad", silence_duration_ms: 500, remove_filler_words: true },
input_audio_noise_reduction: { type: "azure_deep_noise_suppression" },
input_audio_echo_cancellation: { type: "server_echo_cancellation" },
voice: { name: "en-US-Ava:DragonHDLatestNeural", type: "azure-standard" },
},
}));
});
export function sendAudio(pcm16) { // mono PCM16 at 24 kHz (or 16 kHz if configured)
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
// Voice Live follows the Azure OpenAI realtime event set; accept both naming styles
if (ev.type === "response.audio.delta" || ev.type === "response.output_audio.delta")
playPcm16(Buffer.from(ev.delta, "base64"));
if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
if (ev.type === "error") console.error(ev.error);
});
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
You pay for speech twice in cascaded mode
With text LLMs (gpt-5.x, gpt-4.1) you pay LLM text tokens plus Azure speech input and output audio tokens ($15 to $17 in, $25 to $31 out per 1M on standard voices, more for custom). Microsoft does not publish the tokens-per-second rate for Azure speech meters, so measure usage events in a pilot before quoting customers.
Price tier follows the model
You do not choose Pro, Standard or Lite; the model name decides. Mixing a Lite model with a custom voice bills the custom voice at the Pro rate.
Event names differ from OpenAI GA
Voice Live follows the Azure OpenAI realtime reference, whose examples still use flat fields (input_audio_format, voice string or object) and response.audio.delta. Do not paste OpenAI GA session.audio.* configs without checking the Voice Live reference.
Echo cancellation timing
Default server echo cancellation assumes you play audio as soon as it arrives; more than 2 s playback delay degrades it. Use Live-Reference AEC (API 2026-07-15+) if you buffer, mix or resample audio on the client.
Low default TPM
100k tokens per minute per resource is shared by every session on that resource. Long calls re-send history, so request an increase before load testing.
No direct SIP
Phone calls need Azure Communication Services (or another media bridge) in front of Voice Live.
Limited-access features
Custom voice and custom avatar need an approved intake form; plan weeks of lead time.
Preview pieces
azure-realtime, phi4-mm-realtime, MAI Transcribe, MAI-Voice-2-Flash, VASA-2 photo avatars and BYOM are preview and can change.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- 100,000 tokens per minute per resource by default (increase on request)
- Max session duration and concurrency are not stated in Voice Live docs
- Phrase list under 500 words/phrases
- SIP is not supported directly; use Azure Communication Services for telephony
- HD voices only in southeastasia, centralindia, swedencentral, westeurope, eastus, eastus2, westus2 (Voice Live routes synthesis to a supported region)
- Content filtering cannot be modified or disabled (use BYOM for custom filtering)
Models and products
| Name | Status |
|---|---|
| gpt-realtime-2.1 / -datazone / -regional | GA |
| gpt-realtime-2.1-mini, gpt-realtime-mini | GA |
| gpt-realtime, gpt-realtime-1.5 (+ datazone/regional) | GA |
| azure-realtime | Preview |
| phi4-mm-realtime | Preview |
| gpt-5.6-terra, gpt-5.4, gpt-5.2, gpt-5.1, gpt-5, gpt-4.1, gpt-4o | GA |
| gpt-5.6-luna, gpt-5-mini, gpt-4.1-mini, gpt-4o-mini | GA |
| gpt-5-nano, gpt-4.1-nano | GA |
| BYOM (gpt-5.5, gpt-5.4-mini, gpt-5.4-nano and other Foundry deployments) | Preview |
Docs and sources
Docs
Sources used
- learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live
- learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-how-to
- learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-faq
- prices.azure.com/api/retail/prices
Session max duration and concurrency; tokens-per-second for Azure speech input/output meters; which Pro text-output meter ($16 vs $24) applies to which model (likely $24 for gpt-realtime-2.1); retail prices taken from eastus2 meters and may differ by region; WebRTC endpoint details for Voice Live.