Gladia Live (Solaria-1)
Two-step live API: POST to create a session, then stream audio to the returned WebSocket URL. Live uses Solaria-1 with 100+ languages and code-switching.
Overview
Best for: Multilingual live transcription with code-switching and EU vendor preference, with optional realtime translation.
At a glance
100+ languages (Solaria-1). Free credit is EUR 50, one-time. $0.75/hr is the Starter PAYG price; Growth as low as $0.25/hr with commitment. Diarization documented for pre-recorded only. Audio and transcripts retained 3 weeks by default on paid plans.
Declare encoding (e.g. wav/pcm), sample_rate, bit_depth and channels at session init; they must match the chunks you send. Multi-channel supported.
n/a
100+ with code-switching (Solaria-1).
No vendor latency number captured in this pass.
Not re-verified in this pass.
Default retention: paid 3 weeks for audio/transcripts, 1 year metadata; Enterprise custom or zero retention. Certifications not re-verified.
Features
- partial transcripts
- endpointing (silence threshold and max utterance duration)
- code-switching
- custom vocabulary
- realtime translation, summarization, NER and other audio intelligence add-ons
- multi-channel tagging
- webhooks/callbacks
- recorded audio retrievable after the session
Pricing
| What | Price | Unit |
|---|---|---|
| Real-time, Starter (pay as you go) | $0.75 | per hour |
| Real-time, Growth | as low as $0.25 | per hour |
| Enterprise | custom |
$0.25/hr Growth floor to $0.75/hr Starter.
Free tier: One-time EUR 50 credit (vendor estimates 60+ realtime hours); does not renew.
Source: gladia.io
Setup
- Get an API key (x-gladia-key).
- POST https://api.gladia.io/v2/live with encoding, sample_rate, bit_depth, channels and options; the response returns a session id and a WebSocket url.
- Connect to that url (the token is embedded) and send binary audio chunks.
- Read transcript messages (is_final flags partial vs final), send {"type":"stop_recording"} to end, then GET /v2/live/{id} for the full result.
Endpoint
POST https://api.gladia.io/v2/live, then the wss url returned in the response
Authentication
x-gladia-key header on the init POST only; the returned WebSocket URL carries its own session token, so browsers can connect without the API key.
Quick start python
import asyncio, json, os, requests, websockets # pip install requests websockets>=14
cfg = {"encoding": "wav/pcm", "sample_rate": 16000, "bit_depth": 16, "channels": 1,
"messages_config": {"receive_partial_transcripts": True}}
r = requests.post("https://api.gladia.io/v2/live", json=cfg,
headers={"x-gladia-key": os.environ["GLADIA_API_KEY"]})
r.raise_for_status()
ws_url = r.json()["url"] # session-scoped, safe to hand to a browser
async def main():
async with websockets.connect(ws_url) as ws:
async def send():
with open("audio_16k_mono.raw", "rb") as f:
while chunk := f.read(3200): # 100 ms
await ws.send(chunk)
await asyncio.sleep(0.1)
await ws.send(json.dumps({"type": "stop_recording"}))
asyncio.create_task(send())
async for msg in ws:
d = json.loads(msg)
if d.get("type") == "transcript":
kind = "FINAL" if d["data"]["is_final"] else "partial"
print(kind, d["data"]["utterance"]["text"])
asyncio.run(main())
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Audio and transcripts are stored by default
Paid accounts keep audio input and transcripts for 3 weeks by default; free accounts keep everything for 1 year. Zero data retention is Enterprise-only and disables upload and polling endpoints.
Free tier allows only 1 live session
Free accounts get 1 concurrent live session (paid default 30), so any parallel testing needs a paid plan.
Diarization documented for pre-recorded
The speaker diarization page is for pre-recorded audio. Do not assume live speaker labels; use multi-channel input to separate parties on calls.
3-hour session cap
Live sessions are terminated after 3 hours. Long events need a new session and stitching of transcripts.
Format declared up front
encoding, sample_rate, bit_depth and channels must exactly match what you send or the transcript degrades or fails; there is no auto-detection for raw PCM.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- Concurrent live sessions: Free 1, Paid 30 (default), Enterprise on demand; HTTP 429 when exhausted.
- A single live session cannot exceed 3 hours.
- Paid accounts can also queue up to 300 async jobs (not relevant to live).
Models and products
| Name | Status |
|---|---|
| solaria-1 | GA (default, only live model) |
| solaria-3 | GA (pre-recorded only) |
Docs and sources
Docs
Sources used
- gladia.io/pricing
- docs.gladia.io/chapters/introduction/models.md
- docs.gladia.io/chapters/limits-and-specifications/concurrency.md
- docs.gladia.io/chapters/limits-and-specifications/data-retention.md
- docs.gladia.io/chapters/live-stt/quickstart.md
Exact WebSocket message field names in the snippet (data.is_final, data.utterance.text, stop_recording) are from the v2 API as previously documented and were not re-read line by line. Region and latency info not captured.