GA fal

fal Realtime endpoints

WebSocket realtime inference on fal: a persistent connection to a warm runner for fast image-to-image loops (e.g. LCM / SDXL Turbo), plus serverless GPUs where you can deploy your own WebRTC realtime video or world model app.

Est. per minute$0.04 - 0.075

Overview

Best for: Prototyping realtime image transformation loops and hosting your own realtime video models on demand GPUs.

At a glance

Session $/min$0.04
Entry plan $/mo$0
WebRTCYes
WebSocketYes
Self-hostNo
Open weightsNo

Price assumes one self-deployed stream per H100 at $2.49/h (list $4.50/h). Hosted realtime endpoints are per-frame image models billed per model.

Audio in

None for image endpoints

Audio out

None

Languages

Text prompts.

Latency

Vendor: realtime requests skip the queue and reuse a warm runner; first connection can still cold start.

Regions

Not captured.

Compliance

Not verified.

Features

  • persistent WebSocket with msgpack
  • token provider with short-lived JWTs
  • proxy URL pattern
  • deploy your own realtime WebRTC apps on serverless GPUs
  • server errors are never billed

Pricing

WhatPriceUnitNotes
Model API realtime endpointsPer model (unit varies)see model pagefast-lcm-diffusion page showed '$0 per compute second' to our fetcher, which looks like a rendering artefact; verify
H100 serverless GPU$4.50 list, as low as $2.49per hourFor your own deployed realtime apps
H200 serverless GPU$6.00 list, as low as $2.99per hour
B200 serverless GPU$7.99 list, as low as $5.49per hour
How the per-minute estimate was worked out

One H100 runner for a self-deployed realtime app: $2.49-$4.50/h divided by 60 (one stream per GPU assumed)

Free tier: Not verified.

Source: fal.ai

Setup

  1. Create a fal API key.
  2. Add a server-side proxy route or token provider so the browser never sees the key.
  3. npm install @fal-ai/client.
  4. Open fal.realtime.connect to a realtime-capable model and send frames/prompts in a loop.
  5. For live video, deploy your own app with a /webrtc endpoint (see the realtime world model example).

Endpoint

fal.realtime.connect(<model>) or wss://ws.fal.run/{model_id}

Authentication

FAL_KEY via server proxy or short-lived JWT token provider

Quick start javascript

import { fal } from "@fal-ai/client";

// Route through your server so the key stays secret
fal.config({ proxyUrl: "/api/fal/proxy" });

const connection = fal.realtime.connect("fal-ai/fast-lcm-diffusion", {
  onResult: (result) => {
    // result shape: check the model's realtime schema
    console.log(result);
  },
  onError: (err) => console.error(err),
});

// Send a new frame or prompt whenever it changes
connection.send({
  prompt: "a watercolor fox",
  image_url: canvas.toDataURL("image/jpeg", 0.7),
  strength: 0.6,
});
// connection.close() when done

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Image loop, not true video

The hosted realtime endpoints are image-to-image models you call per frame; you build the frame loop, throttling and temporal consistency yourself.

Cold starts

The first connection can hit a cold start; keep the socket open and warm the runner before users arrive.

Self-deployed GPUs bill while connected

A custom realtime app holds a GPU for the whole session; one stream per H100 is $2.49-$4.50 per hour.

Realtime prices not on the main pricing page

Per-model realtime prices were not captured; check each model page or the pricing API before budgeting.

Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.

Limits

  • Only models with an explicit /realtime endpoint work with fal.realtime.connect
  • Cold start on first connection
  • Serverless deploy access needs approval

Models and products

NameStatusNotes
fal-ai/fast-lcm-diffusionGASDXL with Latent Consistency Models; realtime endpoint.
fal-ai/fast-turbo-diffusionGAOptimised SDXL Turbo; realtime endpoint.
Custom realtime apps (e.g. Matrix-Game world model demo)BetaDeploy with fal deploy; /webrtc endpoint; fal Serverless deploy access is approved per account.

Docs and sources

Docs

Sources used

Not fully verified

Per-model realtime prices, result payload shape, free credits, Krea/H3 Max realtime pricing on fal (secondary sources only: $0.08/s after a promo).

Similar avatars + video APIs

Spotted a wrong price or a dead link?