Simli
Low-latency audio-to-video face API: you stream your agent's audio in and get a lip-synced face back over WebRTC or LiveKit. Also offers a managed 'Simli Auto' agent.
Overview
Best for: Developers who already have a voice agent and want the cheapest, fastest face layer.
At a glance
Latency is the vendor claim for the speech-to-video stage only. Price about $0.05/min is from third-party directories. Free: $10 signup plus 50 min/month. maxSessionLength default 60 min (configurable).
PCM16 audio from your TTS (audioInputFormat pcm16)
None added; your audio is passed through with the video
Language-agnostic lip sync from audio (not explicitly verified).
Vendor claim: under 300 ms for the speech-to-video stage.
Not published.
Not verified.
Features
- lip sync from any audio
- bring your own LLM and TTS
- custom face from a photo
- LiveKit and Pipecat plugins
- handleSilence idle animation
- active session count endpoint
Pricing
| What | Price | Unit |
|---|---|---|
| Free plan | $10 credit + 50 min/month | signup credit plus monthly top-up |
| Pay-as-you-go | about $0.05 | per minute |
Third-party reported PAYG rate; unverified
Free tier: $10 on signup and a monthly top-up of 50 minutes (homepage).
Source: simli.com
Setup
- Get an API key from the Simli dashboard.
- Pick a preset face or generate a Trinity face from an image.
- Backend: POST /compose/token to get a session token (set short maxSessionLength and maxIdleTime).
- Frontend or agent: connect via WebRTC SDK or the LiveKit plugin and stream PCM16 audio from your TTS.
Endpoint
POST https://api.simli.ai/compose/token
Authentication
x-simli-api-key header
Quick start javascript
// Server-side: get a short-lived session token
const res = await fetch("https://api.simli.ai/compose/token", {
method: "POST",
headers: {
"x-simli-api-key": process.env.SIMLI_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
faceId: process.env.SIMLI_FACE_ID,
handleSilence: true,
maxSessionLength: 600, // default 3600
maxIdleTime: 60, // default 300
audioInputFormat: "pcm16",
}),
});
const { session_token } = await res.json();
// Give session_token to the Simli client SDK (WebRTC) or use the
// LiveKit plugin, then stream PCM16 audio from your TTS into it.
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
No official price list found
simli.com shows the free tier but no per-minute rate; the ~$0.05/min figure is from third-party directories. Confirm in the dashboard before committing.
Defaults allow 1 hour sessions and 5 minute idle
maxSessionLength defaults to 3600 s and maxIdleTime to 300 s; lower both so abandoned tabs stop billing.
You own the rest of the pipeline
Simli only renders the face; STT, LLM, TTS and the room are separate bills and separate latency.
Two face model families
Trinity and Legacy faces use different creation endpoints; make sure your faceId matches the model you intend.
Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.
Limits
- maxSessionLength default 3600 s
- maxIdleTime default 300 s
- Concurrency per plan not published (Active Session Count endpoint exists)
Models and products
| Name | Status |
|---|---|
| Trinity | GA |
| Legacy faces | GA |
| Gaussian / new facial model | Preview |
Docs and sources
Docs
Sources used
Per-minute price, paid plans, concurrency, Trinity-specific pricing.