GA Microsoft

Azure Speech Voice Live API

Managed voice-agent API in Microsoft Foundry that wraps either native speech-to-speech models (gpt-realtime family, azure-realtime, phi4-mm-realtime) or text LLMs with Azure speech-to-text and 600+ Azure TTS voices, adding noise suppression, server echo cancellation, semantic VAD and avatars. Suits contact centers and branded agents on Azure.

Est. per minute$0.017 - 0.76
1 high-severity warning

Overview

Best for: Contact centers and branded assistants on Azure that want noise suppression, server echo cancellation, multilingual semantic VAD, hundreds of voices, custom voice or a talking avatar without stitching services together.

At a glance

Audio in $/1M tok$32
Audio out $/1M tok$64
Free tierNo
Native S2SYes
ToolsYes
Image inYes
Own LLMYes
Voices600
CloningYes
Languages146
WebRTCYes
WebSocketYes
Phone / SIPNo
HIPAAYes
EU dataYes
Open weightsNo

Supports both native speech-to-speech models and a cascaded STT+LLM+TTS mode. 146 input / 151 output locales. 600+ Azure neural voices; custom voice is limited access. Prices are Pro tier with gpt-realtime-2.1 (Lite from $4/$12). HIPAA via Microsoft BAA for in-scope services; confirm Voice Live scope. Telephony via Azure Communication Services, not direct SIP. BYO model is preview.

Audio in

PCM16 mono at 24 kHz (default) or 16 kHz via input_audio_sampling_rate; Live-Reference AEC mode takes interleaved stereo PCM16 (mic + playback reference). Fixed for the session.

Audio out

PCM16 audio from the native model or Azure TTS; word timestamps and visemes available with Azure voices; avatar video over WebRTC (H.264).

Languages

146 input locales and 151 output locales per the FAQ (Azure speech); azure_semantic_vad_multilingual covers English, Spanish, French, Italian, German, Japanese, Portuguese, Chinese, Korean, Hindi.

Voices

600+ Azure neural voices across 150+ locales, 30+ HD (DragonHD) voices, MAI-Voice-2-Flash (preview), 34 azure-realtime-native voices, native gpt-realtime voices, and custom voice (limited access).

Latency

No numeric vendor claim found; marketed as low-latency. Text-LLM cascades are described by Microsoft as having slightly higher latency than native realtime models.

Regions

10+ Azure regions; LLM availability and processing scope (global, data zone, regional) depend on the resource region and the model suffix.

Compliance

Azure AI services enterprise terms; content filtering always on; data zone and regional model variants for residency. HIPAA BAA via Microsoft for in-scope Azure services (confirm Voice Live scope with Microsoft).

Features

  • azure_semantic_vad and azure_semantic_vad_multilingual turn detection (with filler-word removal), plus server_vad and semantic_vad
  • azure_deep_noise_suppression
  • server_echo_cancellation, plus Live-Reference AEC (client-supplied playback reference, API 2026-07-15+)
  • function calling incl. asynchronous calls; MCP in model mode (not with phi models)
  • Foundry Agent Service integration (agent_id + project_id)
  • phrase lists and custom speech models for recognition
  • custom voice, HD voice temperature, speaking rate, custom lexicon
  • word-level audio timestamps and visemes
  • text-to-speech avatars, photo avatars (VASA-1, VASA-2 preview), custom avatars
  • transcription models: azure-speech, mai-transcribe-2 (preview), whisper-1, gpt-4o-transcribe(-diarize)
  • telephony through Azure Communication Services

Pricing

WhatPriceUnitNotes
Pro: LLM audio input / output (e.g. gpt-realtime-2.1 native audio)$32.00 / $64.00per 1M tokenscached audio $0.40
Pro: LLM text input / output$4.00 / $16.00 (a second Pro text-output meter is $24.00)per 1M tokenscached $0.40; image input $5.00, cached $0.50
Pro: Azure standard speech audio input / output$17.00 / $31.00per 1M tokenscached $0.40
Pro: Azure custom speech audio input / output$40.00 / $55.00per 1M tokenscached $0.40
Standard: LLM audio input / output (e.g. gpt-realtime-2.1-mini)$11.00 / $22.00per 1M tokenscached $0.33
Standard: LLM text input / output$0.66 / $2.64per 1M tokenscached $0.33
Standard: Azure standard speech audio input / output$15.00 / $26.00per 1M tokenscached $0.33
Standard: Azure custom speech audio input / output$39.00 / $50.00per 1M tokens
Lite: LLM audio input / output (phi4-mm-realtime)$4.00 / $12.00per 1M tokenscached $0.04
Lite: LLM text input / output$0.11 / $0.44per 1M tokenscached $0.04
Lite: Azure standard speech audio input / output$15.00 / $25.00per 1M tokens
Lite: Azure custom speech audio input / output$38.00 / $50.00per 1M tokens
BYO model: standard speech audio input / output$12.50 / $23.00per 1M tokenscustom speech $36.00 / $47.00
Avatars, custom voice training and hostingseparateSpeech service pricinge.g. TTS standard avatar realtime $0.50/min, HD standard avatar $0.70/min (Retail Prices API meters)
How the per-minute estimate was worked out

Low = Lite tier phi4-mm-realtime, 1 min user audio (750 tokens x $4/1M) + 1 min model audio (1,200 tokens x $12/1M), single turn. Pro tier with gpt-realtime-2.1 native audio matches Azure OpenAI Global: $0.096 single turn. High = Pro gpt-realtime-2.1, High = a 10 minute call with 50 turns (6 s of user audio + 6 s of model audio per turn, 500-token text system prompt), where the full conversation history is re-billed as input on every turn with no cache hits, total divided by 10 minutes. Standard tier gpt-realtime-2.1-mini: $0.033 low, $0.26 high. Avatar minutes and custom voice are extra.

Audio token rate

Microsoft: Azure OpenAI models ~10 tokens/s input audio and ~20 tokens/s output audio; Phi models ~12.5 tokens/s input, ~20 tokens/s output. Token rate for Azure speech (STT/TTS) meters is not published.

Free tier: No dedicated free tier found for Voice Live (Azure Speech F0 is limited to one concurrent request).

Source: learn.microsoft.com

Setup

  1. Create an Azure subscription and a Microsoft Foundry resource (recommended over a plain Speech resource; Speech resources lack Agent Service and BYOM).
  2. Assign Cognitive Services User and Foundry User roles to your identity, or copy the resource key.
  3. Pick a model by name (it sets the Pro / Standard / Lite price); no deployment is needed for natively supported models.
  4. Connect a server WebSocket to the voice-live/realtime endpoint with api-version and model query parameters.
  5. Send session.update with turn detection, noise suppression, echo cancellation and voice, then stream input_audio_buffer.append.
  6. For browsers use the Voice Live WebRTC flow or relay through your backend; for avatars complete the session.avatar.connect SDP exchange.

Endpoint

wss://<foundry-resource>.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10&model=gpt-realtime-2.1 (older resources: <name>.cognitiveservices.azure.com)

Authentication

Recommended Microsoft Entra token (scope https://ai.azure.com/.default) as Authorization: Bearer. API key works as an api-key header (server only) or an api-key query parameter; never use the query-string key in a browser.

Quick start javascript

import WebSocket from "ws"; // server side (API key in header is not possible from a browser)
const RES = process.env.FOUNDRY_RESOURCE; // Foundry resource name
const url = `wss://${RES}.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10&model=gpt-realtime-2.1`;
const ws = new WebSocket(url, { headers: { "api-key": process.env.FOUNDRY_API_KEY } }); // or Authorization: Bearer <Entra token>
ws.on("open", () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      instructions: "You are a friendly support agent. Keep answers short.",
      input_audio_sampling_rate: 24000,
      turn_detection: { type: "azure_semantic_vad", silence_duration_ms: 500, remove_filler_words: true },
      input_audio_noise_reduction: { type: "azure_deep_noise_suppression" },
      input_audio_echo_cancellation: { type: "server_echo_cancellation" },
      voice: { name: "en-US-Ava:DragonHDLatestNeural", type: "azure-standard" },
    },
  }));
});
export function sendAudio(pcm16) { // mono PCM16 at 24 kHz (or 16 kHz if configured)
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16.toString("base64") }));
}
ws.on("message", (raw) => {
  const ev = JSON.parse(raw.toString());
  // Voice Live follows the Azure OpenAI realtime event set; accept both naming styles
  if (ev.type === "response.audio.delta" || ev.type === "response.output_audio.delta")
    playPcm16(Buffer.from(ev.delta, "base64"));
  if (ev.type === "input_audio_buffer.speech_started") stopPlayback();
  if (ev.type === "error") console.error(ev.error);
});

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

You pay for speech twice in cascaded mode

With text LLMs (gpt-5.x, gpt-4.1) you pay LLM text tokens plus Azure speech input and output audio tokens ($15 to $17 in, $25 to $31 out per 1M on standard voices, more for custom). Microsoft does not publish the tokens-per-second rate for Azure speech meters, so measure usage events in a pilot before quoting customers.

Price tier follows the model

You do not choose Pro, Standard or Lite; the model name decides. Mixing a Lite model with a custom voice bills the custom voice at the Pro rate.

Event names differ from OpenAI GA

Voice Live follows the Azure OpenAI realtime reference, whose examples still use flat fields (input_audio_format, voice string or object) and response.audio.delta. Do not paste OpenAI GA session.audio.* configs without checking the Voice Live reference.

Echo cancellation timing

Default server echo cancellation assumes you play audio as soon as it arrives; more than 2 s playback delay degrades it. Use Live-Reference AEC (API 2026-07-15+) if you buffer, mix or resample audio on the client.

Low default TPM

100k tokens per minute per resource is shared by every session on that resource. Long calls re-send history, so request an increase before load testing.

No direct SIP

Phone calls need Azure Communication Services (or another media bridge) in front of Voice Live.

Limited-access features

Custom voice and custom avatar need an approved intake form; plan weeks of lead time.

Preview pieces

azure-realtime, phi4-mm-realtime, MAI Transcribe, MAI-Voice-2-Flash, VASA-2 photo avatars and BYOM are preview and can change.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • 100,000 tokens per minute per resource by default (increase on request)
  • Max session duration and concurrency are not stated in Voice Live docs
  • Phrase list under 500 words/phrases
  • SIP is not supported directly; use Azure Communication Services for telephony
  • HD voices only in southeastasia, centralindia, swedencentral, westeurope, eastus, eastus2, westus2 (Voice Live routes synthesis to a supported region)
  • Content filtering cannot be modified or disabled (use BYOM for custom filtering)

Models and products

NameStatusNotes
gpt-realtime-2.1 / -datazone / -regionalGAPro tier. Native audio, optional Azure TTS or custom voice output.
gpt-realtime-2.1-mini, gpt-realtime-miniGAStandard tier.
gpt-realtime, gpt-realtime-1.5 (+ datazone/regional)GAPro tier.
azure-realtimePreviewMicrosoft's own dedicated realtime model with 34 azure-realtime-native voices (default ava). Requires API version 2026-01-01-preview or later. Pro tier.
phi4-mm-realtimePreviewLite tier. Phi-4 multimodal audio input + Azure TTS output.
gpt-5.6-terra, gpt-5.4, gpt-5.2, gpt-5.1, gpt-5, gpt-4.1, gpt-4oGAPro tier, cascaded: Azure STT -> LLM -> Azure TTS.
gpt-5.6-luna, gpt-5-mini, gpt-4.1-mini, gpt-4o-miniGAStandard tier, cascaded.
gpt-5-nano, gpt-4.1-nanoGALite tier, cascaded.
BYOM (gpt-5.5, gpt-5.4-mini, gpt-5.4-nano and other Foundry deployments)PreviewBring your own Foundry model deployment.

Docs and sources

Docs

Sources used

Not fully verified

Session max duration and concurrency; tokens-per-second for Azure speech input/output meters; which Pro text-output meter ($16 vs $24) applies to which model (likely $24 for gpt-realtime-2.1); retail prices taken from eastus2 meters and may differ by region; WebRTC endpoint details for Voice Live.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?