GA Amazon Web Services

Amazon Nova 2 Sonic / Nova 2.5 Sonic (Bedrock)

Amazon's native speech-to-speech models on Bedrock, streamed over HTTP/2 bidirectional InvokeModelWithBidirectionalStream. Cheap per token, strong for AWS-native contact centers (Amazon Connect), but with an 8 minute connection limit and a low fixed concurrency quota.

Est. per minute$0.022 - 0.12
2 high-severity warnings

Overview

Best for: AWS-native contact centers and phone agents (Amazon Connect, Twilio, Vonage) that want low per-token cost and in-Region processing.

At a glance

Audio in $/1M tok$3
Audio out $/1M tok$12
Free tierNo
Native S2SYes
ToolsYes
Own LLMNo
Voices16
Languages7
Context tokens1,000,000
Max session min8
Concurrency20
WebRTCNo
WebSocketNo
Phone / SIPNo
HIPAAYes
SOC 2Yes
EU dataYes
Open weightsNo

Transport is HTTP/2 bidirectional streaming. 8 min connection limit. 20 concurrent sessions per account per Region, not adjustable (Nova 2 Sonic). Context 1M for Nova 2 Sonic, 256K for Nova 2.5 Sonic. HIPAA and SOC are Bedrock-level; confirm Nova Sonic in BAA scope. EU via eu-north-1. Telephony via Amazon Connect and partners.

Audio in

audio/lpcm 16-bit mono at 8, 16 or 24 kHz, base64 in audioInput events (~32 ms frames streamed in real time).

Audio out

audio/lpcm 16-bit mono at 8, 16 or 24 kHz, base64 audioOutput events.

Languages

English (US, UK, India, Australia), French, Italian, German, Spanish, Portuguese, Hindi, with automatic language detection and switching.

Voices

16 voice ids in the docs (matthew, tiffany, amy, olivia, lupe, carlos, ambre, florian, lennart, beatrice, lorenzo, tina, carolina, leo, kiara, arjun); polyglot voices can speak every supported language; Polly-compatible voices since March 2026.

Latency

Vendor claim: March 2026 refresh cut user-perceived p50 latency by 150 ms; Nova 2.5 Sonic claimed lower latency than Nova 2 Sonic (no absolute numbers published).

Regions

us-east-1 (N. Virginia), us-west-2 (Oregon), eu-north-1 (Stockholm), ap-northeast-1 (Tokyo). Through Amazon Connect also Singapore, London, Seoul, Frankfurt (Nova 2 Sonic).

Compliance

Bedrock data is not used to train models and stays in the chosen Region (in-Region inference only). Bedrock is HIPAA eligible and covered by AWS SOC reports; confirm Nova Sonic is listed in your BAA scope.

Features

  • turn detection with endpointingSensitivity HIGH / MEDIUM / LOW
  • barge-in handling without losing context
  • function calling with asynchronous tool handling (agent keeps talking while tools run)
  • cross-modal input: text messages during a voice session
  • conversation history injection at session start
  • RAG via your own tools (Bedrock Knowledge Bases integration patterns)
  • telephony via Amazon Connect, Twilio, Vonage, AudioCodes; LiveKit and Pipecat integrations
  • Strands Bidi Agents (GA Oct 2026) handles reconnects across the 8 minute limit

Pricing

WhatPriceUnitNotes
Nova 2 Sonic / Nova 2.5 Sonic speech input$0.003per 1K tokens ($3.00 per 1M)us-east-1
Nova 2 Sonic / Nova 2.5 Sonic speech output$0.012per 1K tokens ($12.00 per 1M)us-east-1
Nova 2 Sonic / Nova 2.5 Sonic text input$0.00033per 1K tokens ($0.33 per 1M)system prompt, history, tool results, cross-modal text
Nova 2 Sonic / Nova 2.5 Sonic text output$0.00275per 1K tokens ($2.75 per 1M)transcripts and tool calls
Nova Sonic (v1) speech input / output$0.0034 / $0.0136per 1K tokenstext $0.00006 in / $0.00024 out per 1K
How the per-minute estimate was worked out

UNVERIFIED token rate. Assuming ~25 speech tokens/s (community figure): low = 1 min user speech (1,500 x $3/1M) + 1 min model speech (1,500 x $12/1M) = $0.0225, single turn. High uses the same 10 minute, 50 turn model as other providers and assumes history is re-processed each turn like other S2S APIs (AWS does not document this). Third-party sites quote roughly $0.015/min.

Audio token rate

Not published by AWS. An AWS re:Post community answer says about 25 tokens per second (unverified). Measure with usageEvent speech token counts.

Free tier: None specific to Nova Sonic (AWS Free Tier credits may apply to new accounts).

Source: pricing.us-east-1.amazonaws.com

Setup

  1. Create an AWS account and pick a supported Region (e.g. us-east-1).
  2. In the Bedrock console confirm access to Amazon Nova 2 Sonic or Nova 2.5 Sonic (and accept any model terms).
  3. Create IAM credentials with bedrock:InvokeModelWithBidirectionalStream, or a Bedrock API key (AWS_BEARER_TOKEN_BEDROCK) for quick tests.
  4. Use an SDK with HTTP/2 bidirectional streaming support (AWS SDK for JavaScript v3 with NodeHttp2Handler, Python aws_sdk_bedrock_runtime, Java, .NET); boto3 does not stream bidirectionally.
  5. Send sessionStart, promptStart, system prompt content, then a long-lived AUDIO content block of audioInput events; read audioOutput, textOutput, toolUse and usageEvent events.
  6. Plan reconnects before 8 minutes and request no more than 20 concurrent sessions per Region per account (use more accounts/Regions or Amazon Connect for scale).

Endpoint

https://bedrock-runtime.{region}.amazonaws.com (InvokeModelWithBidirectionalStream over HTTP/2), modelId amazon.nova-2-sonic-v1:0 or amazon.nova-2-5-sonic

Authentication

AWS SigV4 with IAM credentials (or a Bedrock API key). There is no browser token flow: run the stream on your backend and relay audio from the browser or phone over WebSocket/WebRTC.

Quick start javascript

import { BedrockRuntimeClient, InvokeModelWithBidirectionalStreamCommand } from "@aws-sdk/client-bedrock-runtime";
import { NodeHttp2Handler } from "@smithy/node-http-handler";
import { randomUUID } from "node:crypto";

const client = new BedrockRuntimeClient({ region: "us-east-1", requestHandler: new NodeHttp2Handler({ requestTimeout: 300000 }) });
const P = randomUUID(), SYS = randomUUID(), MIC = randomUUID();
const enc = (event) => ({ chunk: { bytes: new TextEncoder().encode(JSON.stringify({ event })) } });
const audioCfg = { mediaType: "audio/lpcm", sampleRateHertz: 16000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64" };

async function* input(micChunks) {
  yield enc({ sessionStart: { inferenceConfiguration: { maxTokens: 1024, topP: 0.9, temperature: 0.7 }, turnDetectionConfiguration: { endpointingSensitivity: "MEDIUM" } } });
  yield enc({ promptStart: { promptName: P, textOutputConfiguration: { mediaType: "text/plain" },
    audioOutputConfiguration: { ...audioCfg, sampleRateHertz: 24000, voiceId: "matthew" } } });
  yield enc({ contentStart: { promptName: P, contentName: SYS, type: "TEXT", interactive: false, role: "SYSTEM", textInputConfiguration: { mediaType: "text/plain" } } });
  yield enc({ textInput: { promptName: P, contentName: SYS, content: "You are a friendly support agent. Keep answers short." } });
  yield enc({ contentEnd: { promptName: P, contentName: SYS } });
  yield enc({ contentStart: { promptName: P, contentName: MIC, type: "AUDIO", interactive: true, role: "USER", audioInputConfiguration: audioCfg } });
  for await (const pcm of micChunks()) // 16 kHz mono PCM16, ~32 ms frames, streamed in real time
    yield enc({ audioInput: { promptName: P, contentName: MIC, content: Buffer.from(pcm).toString("base64") } });
  yield enc({ contentEnd: { promptName: P, contentName: MIC } });
  yield enc({ promptEnd: { promptName: P } });
  yield enc({ sessionEnd: {} });
}

export async function talk(micChunks, play) {
  const res = await client.send(new InvokeModelWithBidirectionalStreamCommand({ modelId: "amazon.nova-2-sonic-v1:0", body: input(micChunks) }));
  for await (const ev of res.body) {
    if (!ev.chunk?.bytes) continue;
    const msg = JSON.parse(new TextDecoder().decode(ev.chunk.bytes)).event;
    if (msg?.audioOutput) play(Buffer.from(msg.audioOutput.content, "base64")); // 24 kHz PCM16
  }
}

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

20 concurrent sessions, not adjustable

The Nova 2 Sonic model card lists 20 concurrent bidirectional sessions per account per Region as a non-adjustable quota. That caps a single-account deployment at 20 simultaneous calls per Region; plan multi-Region, multi-account or Amazon Connect for real call volume.

8 minute connection limit

Each bidirectional stream ends at 8 minutes. You must open a new stream and re-send system prompt and history (paying text input again) at a natural pause, or use Strands Bidi Agents which automates this.

Token rate not published

AWS bills speech tokens but does not publish tokens per second, so per-minute cost can only be measured from usageEvent data. Budget conservatively until you have pilot numbers.

Strict event ordering

Every event must carry the right promptName and contentName, history must come before audio, and sessions must close with contentEnd, promptEnd, sessionEnd. Mistakes cause opaque stream errors or orphaned sessions.

SDK support is uneven

Bidirectional streaming needs HTTP/2 SDK support; boto3 does not support it, and the Python path uses a separate experimental SDK. Keep audio frames flowing in real time or the stream can time out.

Silent in-place model refreshes

Nova 2 Sonic was updated in place in March and May 2026 with no API change. Behaviour (voices, turn-taking) can shift under a fixed model id; keep regression tests for prompts and tools.

Limited languages

Seven languages; no Japanese, Chinese or Arabic despite the Tokyo Region. Check before selling into those markets.

Bedrock features missing

Guardrails, Knowledge Bases, Agents, prompt management and token counting are not supported on the Nova 2 Sonic bidirectional endpoint; implement safety and retrieval in your tools.

Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.

Limits

  • Connection limit 8 minutes; renew the connection and carry history forward (AWS samples and Strands Bidi Agents do this)
  • Default quota 20 concurrent InvokeModelWithBidirectionalStream sessions per account per Region for Nova 2 Sonic, documented as not adjustable
  • Context window: Nova 2 Sonic 1M tokens (64K output); Nova 2.5 Sonic 256K per launch post
  • Conversation history can only be injected once, after the system prompt and before audio streaming
  • In-Region inference only (no geo or global cross-region profiles)
  • Standard service tier only (no Priority, Flex or Reserved)

Models and products

NameStatusNotes
amazon.nova-2-sonic-v1:0GAGA Dec 2, 2025; refreshed in place Mar 2026 (p50 latency -150 ms, Polly-compatible voices, better 8 kHz turn-taking) and May 2026 (88 percent fewer speech hallucinations on AWS internal set). 1M token context, 64K max output. EOL no sooner than Dec 2, 2026.
amazon.nova-2-5-sonicGAGA Oct 5, 2026: better reasoning, instruction following and tool calling, lower latency, 256K context, same price as Nova 2 Sonic. Model id taken from the Strands Bidi Agents launch post; confirm in the Bedrock console.
amazon.nova-sonic-v1:0GAOriginal Nova Sonic (Apr 2025); 300K context, English-focused; slightly higher speech prices. Superseded.

Docs and sources

Docs

Sources used

Not fully verified

Speech tokens per second (only a community figure of ~25/s); whether history is re-billed per turn; official Nova 2.5 Sonic model id (amazon.nova-2-5-sonic from the Strands blog), its concurrency quota and model card; whether the 20-session quota also applies to Nova 2.5 Sonic; prices in Regions other than us-east-1.

Similar voice-to-voice APIs

Spotted a wrong price or a dead link?