GA LiveKit

LiveKit Agents + LiveKit Cloud

Open-source (Apache-2.0) Python/Node framework for realtime voice agents on top of LiveKit's WebRTC SFU, with a managed cloud that hosts agents, provides telephony and a pay-per-use inference gateway for STT, LLM and TTS.

Est. per minute$0.03 - 0.08

Overview

Best for: Developers who want full code control of the voice pipeline with a managed WebRTC/telephony layer, or a path to self-hosting.

At a glance

Platform $/min$0.01
All-in priceNo
Entry plan $/mo$50
Free tierYes
Free credit $$2.5
Phone numbersYes
Own LLMYes
Own keysYes
RecordingYes
Noise cancelYes
Smart turnsYes
Concurrency20
WebRTCYes
Phone / SIPYes
SOC 2Yes
EU dataYes
Self-hostYes
Open weightsYes

Free Build plan: 1,000 agent minutes, 5 concurrent, $2.50 inference credit. $0.01/min applies after included minutes. Ship concurrency 20; Scale listed up to 600 (ambiguous). HIPAA/BAA not confirmed.

Audio in

Opus over WebRTC; SIP/PSTN calls bridged in

Audio out

Opus over WebRTC; telephony codecs on SIP legs

Languages

Depends on the chosen STT/TTS/LLM; LiveKit Inference restricts to EU-hosted models for EU agents.

Latency

No single vendor number verified; depends on models and region.

Regions

Global SFU; project data region US (default) or EU (Frankfurt), fixed at creation; agents deployable to regions such as eu-central.

Compliance

SOC 2 Type II (Security, Availability, Confidentiality) and GDPR per LiveKit security page; HIPAA/BAA not verified.

Features

  • turn detection (transformer model + inference.TurnDetector)
  • noise cancellation (ai-coustics, BVC)
  • telephony (LiveKit Phone Numbers, third-party SIP, Twilio connector)
  • multi-provider plug-ins
  • recording
  • observability/traces
  • MCP tools
  • realtime speech-to-speech models
  • agent hosting (lk agent create)

Pricing

WhatPriceUnitNotes
Agent session minutes$0.01per minuteAfter included minutes (1,000 Build, 5,000 Ship, 50,000 Scale). Hosting/orchestration only; excludes model costs.
Agent session recordings$0.005per minuteAfter included minutes.
WebRTC participant minutes$0.0005 (Ship) / $0.0004 (Scale)per minute5,000 / 150,000 / 1.5M included on Build / Ship / Scale.
Third-party SIP minutes$0.004 (Ship) / $0.003 (Scale)per minuteYour own SIP trunk, carrier charges billed by your carrier.
US local inbound (LiveKit number)$0.01per minutePlus $1.00/month per extra number; toll-free $2.00/month and $0.02/min inbound.
Downstream data$0.12 (Ship) / $0.10 (Scale)per GBAfter 50GB/250GB/3TB.
Plans$0 / $50 / $500 per monthper monthBuild / Ship / Scale.
Inference STT/LLM/TTSe.g. STT $0.0025-$0.0048, TTS $0.018-$0.03, LLM GPT-4o mini $0.0006per minuteOptional; or bring your own provider keys via plugins.
How the per-minute estimate was worked out

LiveKit's own calculator default is $0.0479/min (agent session $0.01 + telephony $0.01 + STT $0.0075 + LLM $0.0014 + TTS $0.009 + observability $0.01). Low end: WebRTC only, cheap STT/TTS, small LLM. High end: phone calls, premium TTS (Cartesia/ElevenLabs class) and a larger LLM. Speech-to-speech realtime models cost much more and are priced in the model segment.

Free tier: Build plan: 1,000 agent session minutes/month, $2.50 inference credit, 1 US number with 50 inbound minutes, no credit card.

Source: livekit.com

Setup

  1. Install the LiveKit CLI and run 'lk cloud auth' to link a LiveKit Cloud project.
  2. 'lk agent init my-agent --template agent-starter-python' creates the project and .env.local with credentials.
  3. 'uv sync', then 'lk agent dev' to run locally and talk to it from the Agents Playground or console mode.
  4. 'lk agent create' to deploy the agent to LiveKit Cloud; attach a phone number or SIP trunk with a dispatch rule for phone calls.

Endpoint

wss://<project>.livekit.cloud

Authentication

LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET (JWT access tokens for clients)

Quick start python

# agent.py - shape per LiveKit voice AI quickstart (Oct 2026); model ids from the docs example
from dotenv import load_dotenv
from livekit import agents
from livekit.agents import AgentServer, AgentSession, Agent, inference, TurnHandlingOptions

load_dotenv(".env.local")

class Assistant(Agent):
    def __init__(self):
        super().__init__(instructions="You are a concise, friendly voice assistant.")

server = AgentServer()

@server.rtc_session(agent_name="my-agent")
async def my_agent(ctx: agents.JobContext):
    session = AgentSession(
        stt=inference.STT(model="assemblyai/universal-3-6-pro", language="en"),
        llm=inference.LLM(model="google/gemma-4-31b-it"),
        tts=inference.TTS(model="fishaudio/s2.1-pro", voice="<voice-id>"),
        turn_handling=TurnHandlingOptions(turn_detection=inference.TurnDetector()),
    )
    await session.start(room=ctx.room, agent=Assistant())
    await session.generate_reply(instructions="Greet the user and offer your assistance.")

if __name__ == "__main__":
    agents.cli.run_app(server)  # python agent.py dev

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Agent session minutes are only the hosting fee

$0.01/min covers running your agent; STT, LLM and TTS are billed separately (via LiveKit Inference credits or your own provider keys), and phone minutes are billed again.

Free inference credit is tiny

Build includes $2.50 of inference (~50 minutes). After that, calls fail or need your own provider keys unless you upgrade.

Data region is immutable

Project data region (US or EU) is chosen at project creation and cannot be changed. Telephony, storage and third-party plugins each need their own regional configuration for GDPR setups.

API churn between versions

The quickstart moved to AgentServer/@server.rtc_session and TurnHandlingOptions; older WorkerOptions/entrypoint examples on blogs may not match the current SDK. Pin versions.

Turn detector model license

The Agents framework is Apache-2.0, but LiveKit's turn detection models are under the separate LiveKit Model License. Read it before self-hosting commercially.

HIPAA status not confirmed

Official security page lists SOC 2 Type II and GDPR; a HIPAA/BAA statement was not confirmed in this research. Ask sales before handling PHI.

Plus 12 warnings that apply to all platforms and telephony APIs. See category warnings.

Limits

  • Concurrent agent sessions: 5 Build, 20 Ship, up to 600 Scale (page ambiguous), custom Enterprise
  • Inference concurrency: 5 / 20 / 50
  • Agent deployments: 1 / 2 / 4
  • WebRTC concurrent connections: 100 / 1,000 / 5,000

Models and products

NameStatusNotes
Build planGA$0/mo, no card: 1,000 agent session minutes, 5 concurrent agent sessions, 1 deployment, $2.50 inference credit (~50 min), 1 free US local number with 50 inbound minutes.
Ship planGA$50/mo: 5,000 agent session minutes then $0.01/min, 20 concurrent sessions, $5 inference credit.
Scale planGA$500/mo: 50,000 agent session minutes then $0.01/min, inference concurrency 50, $50 inference credit and discounted model prices. Pricing page lists 'up to 600' concurrent sessions but the table layout is ambiguous.
EnterpriseGACustom.
LiveKit Inference (gateway)GAOne API key for many STT/LLM/TTS models, billed per minute on the LiveKit invoice (e.g. Deepgram Nova-3 $0.0048/min, AssemblyAI Universal-Streaming $0.0025/min, Cartesia Sonic 3 $0.03/min, Deepgram Aura-2 $0.018/min on Build/Ship). Some TTS models are listed as free.
Self-hosted LiveKit server + AgentsGAOpen source; you run the SFU, TURN, agent workers and pay providers directly.

Docs and sources

Docs

Sources used

Not fully verified

Scale plan concurrency (table on pricing page is ambiguous); HIPAA/BAA availability; exact model ids change frequently.

Similar platforms APIs

Spotted a wrong price or a dead link?