LiveKit Agents + LiveKit Cloud
Open-source (Apache-2.0) Python/Node framework for realtime voice agents on top of LiveKit's WebRTC SFU, with a managed cloud that hosts agents, provides telephony and a pay-per-use inference gateway for STT, LLM and TTS.
Overview
Best for: Developers who want full code control of the voice pipeline with a managed WebRTC/telephony layer, or a path to self-hosting.
At a glance
Free Build plan: 1,000 agent minutes, 5 concurrent, $2.50 inference credit. $0.01/min applies after included minutes. Ship concurrency 20; Scale listed up to 600 (ambiguous). HIPAA/BAA not confirmed.
Opus over WebRTC; SIP/PSTN calls bridged in
Opus over WebRTC; telephony codecs on SIP legs
Depends on the chosen STT/TTS/LLM; LiveKit Inference restricts to EU-hosted models for EU agents.
No single vendor number verified; depends on models and region.
Global SFU; project data region US (default) or EU (Frankfurt), fixed at creation; agents deployable to regions such as eu-central.
SOC 2 Type II (Security, Availability, Confidentiality) and GDPR per LiveKit security page; HIPAA/BAA not verified.
Features
- turn detection (transformer model + inference.TurnDetector)
- noise cancellation (ai-coustics, BVC)
- telephony (LiveKit Phone Numbers, third-party SIP, Twilio connector)
- multi-provider plug-ins
- recording
- observability/traces
- MCP tools
- realtime speech-to-speech models
- agent hosting (lk agent create)
Pricing
| What | Price | Unit |
|---|---|---|
| Agent session minutes | $0.01 | per minute |
| Agent session recordings | $0.005 | per minute |
| WebRTC participant minutes | $0.0005 (Ship) / $0.0004 (Scale) | per minute |
| Third-party SIP minutes | $0.004 (Ship) / $0.003 (Scale) | per minute |
| US local inbound (LiveKit number) | $0.01 | per minute |
| Downstream data | $0.12 (Ship) / $0.10 (Scale) | per GB |
| Plans | $0 / $50 / $500 per month | per month |
| Inference STT/LLM/TTS | e.g. STT $0.0025-$0.0048, TTS $0.018-$0.03, LLM GPT-4o mini $0.0006 | per minute |
LiveKit's own calculator default is $0.0479/min (agent session $0.01 + telephony $0.01 + STT $0.0075 + LLM $0.0014 + TTS $0.009 + observability $0.01). Low end: WebRTC only, cheap STT/TTS, small LLM. High end: phone calls, premium TTS (Cartesia/ElevenLabs class) and a larger LLM. Speech-to-speech realtime models cost much more and are priced in the model segment.
Free tier: Build plan: 1,000 agent session minutes/month, $2.50 inference credit, 1 US number with 50 inbound minutes, no credit card.
Source: livekit.com
Setup
- Install the LiveKit CLI and run 'lk cloud auth' to link a LiveKit Cloud project.
- 'lk agent init my-agent --template agent-starter-python' creates the project and .env.local with credentials.
- 'uv sync', then 'lk agent dev' to run locally and talk to it from the Agents Playground or console mode.
- 'lk agent create' to deploy the agent to LiveKit Cloud; attach a phone number or SIP trunk with a dispatch rule for phone calls.
Endpoint
wss://<project>.livekit.cloud
Authentication
LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET (JWT access tokens for clients)
Quick start python
# agent.py - shape per LiveKit voice AI quickstart (Oct 2026); model ids from the docs example
from dotenv import load_dotenv
from livekit import agents
from livekit.agents import AgentServer, AgentSession, Agent, inference, TurnHandlingOptions
load_dotenv(".env.local")
class Assistant(Agent):
def __init__(self):
super().__init__(instructions="You are a concise, friendly voice assistant.")
server = AgentServer()
@server.rtc_session(agent_name="my-agent")
async def my_agent(ctx: agents.JobContext):
session = AgentSession(
stt=inference.STT(model="assemblyai/universal-3-6-pro", language="en"),
llm=inference.LLM(model="google/gemma-4-31b-it"),
tts=inference.TTS(model="fishaudio/s2.1-pro", voice="<voice-id>"),
turn_handling=TurnHandlingOptions(turn_detection=inference.TurnDetector()),
)
await session.start(room=ctx.room, agent=Assistant())
await session.generate_reply(instructions="Greet the user and offer your assistance.")
if __name__ == "__main__":
agents.cli.run_app(server) # python agent.py dev
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Agent session minutes are only the hosting fee
$0.01/min covers running your agent; STT, LLM and TTS are billed separately (via LiveKit Inference credits or your own provider keys), and phone minutes are billed again.
Free inference credit is tiny
Build includes $2.50 of inference (~50 minutes). After that, calls fail or need your own provider keys unless you upgrade.
Data region is immutable
Project data region (US or EU) is chosen at project creation and cannot be changed. Telephony, storage and third-party plugins each need their own regional configuration for GDPR setups.
API churn between versions
The quickstart moved to AgentServer/@server.rtc_session and TurnHandlingOptions; older WorkerOptions/entrypoint examples on blogs may not match the current SDK. Pin versions.
Turn detector model license
The Agents framework is Apache-2.0, but LiveKit's turn detection models are under the separate LiveKit Model License. Read it before self-hosting commercially.
HIPAA status not confirmed
Official security page lists SOC 2 Type II and GDPR; a HIPAA/BAA statement was not confirmed in this research. Ask sales before handling PHI.
Plus 12 warnings that apply to all platforms and telephony APIs. See category warnings.
Limits
- Concurrent agent sessions: 5 Build, 20 Ship, up to 600 Scale (page ambiguous), custom Enterprise
- Inference concurrency: 5 / 20 / 50
- Agent deployments: 1 / 2 / 4
- WebRTC concurrent connections: 100 / 1,000 / 5,000
Models and products
| Name | Status |
|---|---|
| Build plan | GA |
| Ship plan | GA |
| Scale plan | GA |
| Enterprise | GA |
| LiveKit Inference (gateway) | GA |
| Self-hosted LiveKit server + Agents | GA |
Docs and sources
Docs
Sources used
- livekit.com/pricing
- docs.livekit.io/agents/start/voice-ai-quickstart/
- github.com/livekit/agents
- livekit.com/security/overview
- docs.livekit.io/deploy/admin/regions/data-residency
Scale plan concurrency (table on pricing page is ambiguous); HIPAA/BAA availability; exact model ids change frequently.