fal Realtime endpoints
WebSocket realtime inference on fal: a persistent connection to a warm runner for fast image-to-image loops (e.g. LCM / SDXL Turbo), plus serverless GPUs where you can deploy your own WebRTC realtime video or world model app.
Overview
Best for: Prototyping realtime image transformation loops and hosting your own realtime video models on demand GPUs.
At a glance
Price assumes one self-deployed stream per H100 at $2.49/h (list $4.50/h). Hosted realtime endpoints are per-frame image models billed per model.
None for image endpoints
None
Text prompts.
Vendor: realtime requests skip the queue and reuse a warm runner; first connection can still cold start.
Not captured.
Not verified.
Features
- persistent WebSocket with msgpack
- token provider with short-lived JWTs
- proxy URL pattern
- deploy your own realtime WebRTC apps on serverless GPUs
- server errors are never billed
Pricing
| What | Price | Unit |
|---|---|---|
| Model API realtime endpoints | Per model (unit varies) | see model page |
| H100 serverless GPU | $4.50 list, as low as $2.49 | per hour |
| H200 serverless GPU | $6.00 list, as low as $2.99 | per hour |
| B200 serverless GPU | $7.99 list, as low as $5.49 | per hour |
One H100 runner for a self-deployed realtime app: $2.49-$4.50/h divided by 60 (one stream per GPU assumed)
Free tier: Not verified.
Source: fal.ai
Setup
- Create a fal API key.
- Add a server-side proxy route or token provider so the browser never sees the key.
- npm install @fal-ai/client.
- Open fal.realtime.connect to a realtime-capable model and send frames/prompts in a loop.
- For live video, deploy your own app with a /webrtc endpoint (see the realtime world model example).
Endpoint
fal.realtime.connect(<model>) or wss://ws.fal.run/{model_id}
Authentication
FAL_KEY via server proxy or short-lived JWT token provider
Quick start javascript
import { fal } from "@fal-ai/client";
// Route through your server so the key stays secret
fal.config({ proxyUrl: "/api/fal/proxy" });
const connection = fal.realtime.connect("fal-ai/fast-lcm-diffusion", {
onResult: (result) => {
// result shape: check the model's realtime schema
console.log(result);
},
onError: (err) => console.error(err),
});
// Send a new frame or prompt whenever it changes
connection.send({
prompt: "a watercolor fox",
image_url: canvas.toDataURL("image/jpeg", 0.7),
strength: 0.6,
});
// connection.close() when done
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Image loop, not true video
The hosted realtime endpoints are image-to-image models you call per frame; you build the frame loop, throttling and temporal consistency yourself.
Cold starts
The first connection can hit a cold start; keep the socket open and warm the runner before users arrive.
Self-deployed GPUs bill while connected
A custom realtime app holds a GPU for the whole session; one stream per H100 is $2.49-$4.50 per hour.
Realtime prices not on the main pricing page
Per-model realtime prices were not captured; check each model page or the pricing API before budgeting.
Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.
Limits
- Only models with an explicit /realtime endpoint work with fal.realtime.connect
- Cold start on first connection
- Serverless deploy access needs approval
Models and products
| Name | Status |
|---|---|
| fal-ai/fast-lcm-diffusion | GA |
| fal-ai/fast-turbo-diffusion | GA |
| Custom realtime apps (e.g. Matrix-Game world model demo) | Beta |
Docs and sources
Docs
Sources used
- fal.ai/docs/model-apis/real-time
- fal.ai/pricing
- fal.ai/models/fal-ai/fast-lcm-diffusion
- fal.ai/docs/examples/video-generation/deploy-realtime-world-model
Per-model realtime prices, result payload shape, free credits, Krea/H3 Max realtime pricing on fal (secondary sources only: $0.08/s after a promo).