Overshoot
Realtime vision API: publish a live camera to a stream over LiveKit, then ask OpenAI-style chat completions about the latest frames using fast open VLMs it hosts, or pass through to Gemini, Claude or GPT.
Overview
Best for: Cheap, fast 'what is happening on camera right now' questions with open VLMs.
At a glance
Token-billed prepaid credits from $1; no stream-time charge shown. Vendor targets sub-second TTFT on hosted models without a number. Frames kept 600 s.
None
None (text)
Depends on chosen model.
Vendor: hosted models sized for sub-second time-to-first-token on single-frame inputs; no measured figure.
Not published.
Not verified.
Features
- live stream frame references (ovs://streams/{id}?frame_index=-1)
- OpenAI-compatible chat completions
- open VLMs and proprietary passthrough
- public pricing endpoint
- prepaid credits from $1
Pricing
| What | Price | Unit |
|---|---|---|
| gemma-4-26B-A4B (hosted) | $0.06 in / $0.33 out | per 1M tokens |
| Qwen3.6-27B-FP8 (hosted) | $0.29 in / $2.40 out | per 1M tokens |
| gemini-3-flash-preview passthrough | $0.50 in / $3.00 out | per 1M tokens |
| claude-sonnet-4-6 passthrough | $3.00 in / $15.00 out | per 1M tokens |
| Stream time | Not shown |
Token-based; depends on frames per query and tokens per frame (not published)
Free tier: None found; prepaid credits, minimum $1.
Source: api.overshoot.ai
Setup
- Buy prepaid credits and create an API key.
- POST /v1beta/streams to create a stream; get publish.url and publish.token.
- Publish your webcam to that LiveKit room.
- POST /v1beta/chat/completions referencing ovs://streams/{id}?frame_index=-1 for the latest frame.
Endpoint
https://api.overshoot.ai/v1beta
Authentication
Authorization: Bearer <OVERSHOOT_API_KEY>
Quick start javascript
const H = {
Authorization: `Bearer ${process.env.OVERSHOOT_API_KEY}`,
"Content-Type": "application/json",
};
const API = "https://api.overshoot.ai/v1beta";
// 1) create a stream, then publish the camera to publish.url
// with publish.token using the LiveKit client SDK
const stream = await (await fetch(`${API}/streams`, { method: "POST", headers: H })).json();
// 2) ask about the newest frame
const r = await fetch(`${API}/chat/completions`, {
method: "POST",
headers: H,
body: JSON.stringify({
model: "google/gemma-4-26B-A4B-it",
messages: [{ role: "user", content: [
{ type: "text", text: "Is anyone at the door?" },
{ type: "image_url", image_url: { url: `ovs://streams/${stream.id}?frame_index=-1` } },
]}],
}),
});
console.log((await r.json()).choices[0].message.content);
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Young startup, beta API
v1beta paths and a YC W26 company; expect breaking changes and keep an abstraction layer.
Passthrough is not realtime
Claude/GPT/Gemini passthrough adds seconds of latency; use hosted Gemma/Qwen for sub-second loops.
You poll, it does not push
Analysis happens per chat completion you send; continuous monitoring means calling it on a timer, which multiplies token cost.
Short frame history
Only the last 600 seconds of frames are addressable.
Plus 9 warnings that apply to all avatars and live video APIs. See category warnings.
Limits
- Frames retained for 600 seconds
- Frames may be compressed or resized before inference
- A 'ready' model can still return 503
Models and products
| Name | Status |
|---|---|
| google/gemma-4-26B-A4B-it | GA |
| google/gemma-4-31B-it | GA |
| Qwen/Qwen3.6-27B-FP8 | GA |
| Qwen/Qwen3.6-35B-A3B-FP8 | GA |
| Hcompany/Holo3-35B-A3B and Holo-3.1 | GA |
| Passthrough: Gemini 3.x, Claude 4.x, GPT-5.4 family | GA |
Docs and sources
Docs
Sources used
- api.overshoot.ai/billing/pricing
- docs.overshoot.ai/models.md
- docs.overshoot.ai/quickstart.md
- ycombinator.com/companies/overshoot
Stream-time charges, concurrency, tokens per frame, latency.