Vercel AI SDK realtime (via AI Gateway)
Beta realtime voice support in the AI SDK through Vercel AI Gateway: gateway.experimental_realtime mints connection tokens server-side and a useRealtime React hook handles mic capture and playback for speech-to-speech models.
Overview
Best for: Next.js/React apps already using the AI SDK and AI Gateway that want a quick browser voice mode.
At a glance
Beta/canary API. SDK is free; AI Gateway passes through model cost (markup on realtime models unverified).
Browser mic
Browser playback
Per model.
Model-dependent.
Vercel AI Gateway.
Not covered.
Features
- one gateway key for multiple realtime model vendors
- React hook
- batch STT/TTS via AI Gateway
Pricing
| What | Price | Unit |
|---|---|---|
| AI SDK | $0 | license |
Equals the underlying realtime model cost; see models segment.
Free tier: SDK open source; AI Gateway has its own terms.
Source: vercel.com
Setup
- pnpm add ai@canary @ai-sdk/gateway@canary @ai-sdk/react@canary (or AI SDK 7 per changelog).
- Create a server route that uses gateway.experimental_realtime getToken to mint a connection.
- Use the useRealtime hook in a React client for mic and playback.
Endpoint
Vercel AI Gateway
Authentication
AI_GATEWAY_API_KEY or Vercel OIDC
Quick start javascript
// Vercel AI SDK realtime is Beta and its signatures are changing; follow
// https://vercel.com/docs/ai-gateway/modalities/realtime for current code.
// Install: pnpm add ai@canary @ai-sdk/gateway@canary @ai-sdk/react@canary
// Server: gateway.experimental_realtime -> getToken() for a model such as "openai/gpt-realtime-2"
// Client: const rt = useRealtime({ ...token endpoint... }) -> start/stop mic + playback
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Canary/Beta API
Exact function signatures were not stable enough to publish a full snippet; expect breaking changes.
Version confusion
Docs say canary releases; changelog says AI SDK 7. Use whichever the docs page you follow specifies.
Speech-to-speech only models
Some gateway realtime models do not support transcription, so you cannot get user transcripts from them.
Plus 12 warnings that apply to all platforms and telephony APIs. See category warnings.
Limits
- Beta / canary API surface
- Some models are speech-to-speech only (no transcription)
Models and products
| Name | Status |
|---|---|
| gateway.experimental_realtime + useRealtime | Beta |
Docs and sources
Docs
Sources used
- vercel.com/docs/ai-gateway/modalities/realtime
- vercel.com/changelog/realtime-voice-speech-and-transcription-now-supported-on-a...
Exact API signatures; gateway markup on realtime models.