Phonic
A voice-agent platform built around its own speech-to-speech model ('merritt'), with tools, evals, conversation replay, SIP and LiveKit support. Aimed at production phone agents where reliability matters more than lowest price.
Overview
Best for: Production phone agents in regulated settings (HIPAA) wanting a native speech-to-speech model with tooling for evals and replay.
At a glance
Latency is a vendor claim of under 500 ms speech in to speech out. Price is the published starting rate. HIPAA and SOC 2 are vendor claims. LiveKit plugin available.
pcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000
Same options
51 languages (per docs index).
Named preset voices (e.g. sabrina, grant, eleanor, nolan) with accents; listable via voices.list.
Vendor claim: sub-500 ms speech-in to speech-out.
Not stated.
HIPAA and SOC 2 compliance, 99.9% uptime SLA (vendor claim).
Features
- webhook, WebSocket, context, transfer, MCP and built-in tools
- interruption handling
- pronunciation dictionary
- multilingual switching
- push-to-talk
- conversation replay for testing
- evals
- audit logs and retention controls
- background noise option
Pricing
| What | Price | Unit |
|---|---|---|
| Conversation | from $0.15 | per minute |
Published starting price; volume pricing via sales.
Not token-billed.
Free tier: Not stated.
Source: docs.phonic.ai
Setup
- Create a Phonic account and API key; build an agent in the dashboard.
- Connect with the SDK (or raw WebSocket to /v1/sts/ws) using Authorization: Bearer <key>; browsers use session tokens.
- Send a config message naming the agent, then forward audio_chunk messages both ways.
Endpoint
wss://api.phonic.ai/v1/sts/ws
Authentication
Authorization: Bearer <PHONIC_API_KEY>; browser/mobile: ?session_token=... or your own JWTs
Quick start javascript
// npm i phonic (create an agent in the Phonic dashboard first)
import { PhonicClient } from "phonic";
const phonic = new PhonicClient({ apiKey: process.env.PHONIC_API_KEY });
// Opens wss://api.phonic.ai/v1/sts/ws with Authorization: Bearer <key>
const socket = await phonic.conversations.connect();
socket.on("message", (msg) => {
if (msg.type === "audio_chunk") play(msg.audio); // base64 agent audio (+ transcript)
});
await socket.sendConfig({
type: "config",
agent: "my-agent", // agent name from the dashboard
// input_format / output_format: pcm_44100 (default), pcm_24000, pcm_16000, pcm_8000, mulaw_8000
});
// Forward caller/mic audio (base64) as it arrives
export async function sendAudio(b64) {
await socket.sendAudioChunk({ type: "audio_chunk", audio: b64 });
}
function play(b64) {}
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Most expensive per minute here
From $0.15/min is 2-3x Ultravox or Deepgram Standard; check whether the price includes telephony and what volume discounts exist.
No public rate card
Only a 'starting at' price is published; enterprise terms vary.
Single proprietary model
Only 'merritt' is offered; you cannot swap in your own LLM for reasoning-heavy tasks.
Default 44.1 kHz input
Default input format is pcm_44100; telephony bridges must set mulaw_8000 explicitly or audio will be wrong.
Live conversation WebSocket is Preview
The docs label the live conversation WebSocket as Preview; expect protocol changes.
Plus 14 warnings that apply to all voice-to-voice APIs. See category warnings.
Limits
- Rate limits per organization: 500 requests/second shared across /sts/ws and SIP outbound; 5 requests/second for /conversations/outbound_call
Models and products
| Name | Status |
|---|---|
| merritt | GA |
Docs and sources
Docs
Sources used
- docs.phonic.ai/
- docs.phonic.ai/llms.txt
- docs.phonic.ai/api-reference/conversations/conversations.md
- docs.phonic.ai/docs/platform/rate_limits.md
- docs.phonic.ai/docs/build/agents/voices.md
What the $0.15 covers (telephony?), concurrency caps, regions, free trial.