GA Gladia

Gladia Live (Solaria-1)

Two-step live API: POST to create a session, then stream audio to the returned WebSocket URL. Live uses Solaria-1 with 100+ languages and code-switching.

Est. per minute$0.0042 - 0.013
1 high-severity warning

Overview

Best for: Multilingual live transcription with code-switching and EU vendor preference, with optional realtime translation.

At a glance

$/hour$0.75
Free tierYes
Turn detectYes
KeytermsYes
PartialsYes
Mixed langsYes
Languages100
Max session min180
Concurrency30
WebRTCNo
WebSocketYes
gRPCNo
Self-hostNo
Open weightsNo

100+ languages (Solaria-1). Free credit is EUR 50, one-time. $0.75/hr is the Starter PAYG price; Growth as low as $0.25/hr with commitment. Diarization documented for pre-recorded only. Audio and transcripts retained 3 weeks by default on paid plans.

Audio in

Declare encoding (e.g. wav/pcm), sample_rate, bit_depth and channels at session init; they must match the chunks you send. Multi-channel supported.

Audio out

n/a

Languages

100+ with code-switching (Solaria-1).

Latency

No vendor latency number captured in this pass.

Regions

Not re-verified in this pass.

Compliance

Default retention: paid 3 weeks for audio/transcripts, 1 year metadata; Enterprise custom or zero retention. Certifications not re-verified.

Features

  • partial transcripts
  • endpointing (silence threshold and max utterance duration)
  • code-switching
  • custom vocabulary
  • realtime translation, summarization, NER and other audio intelligence add-ons
  • multi-channel tagging
  • webhooks/callbacks
  • recorded audio retrievable after the session

Pricing

WhatPriceUnitNotes
Real-time, Starter (pay as you go)$0.75per hourPricing page lists it as a 'starting at' price.
Real-time, Growthas low as $0.25per hourRequires upfront commitment; exact rate not published.
EnterprisecustomCustom models, zero data retention option.
How the per-minute estimate was worked out

$0.25/hr Growth floor to $0.75/hr Starter.

Free tier: One-time EUR 50 credit (vendor estimates 60+ realtime hours); does not renew.

Source: gladia.io

Setup

  1. Get an API key (x-gladia-key).
  2. POST https://api.gladia.io/v2/live with encoding, sample_rate, bit_depth, channels and options; the response returns a session id and a WebSocket url.
  3. Connect to that url (the token is embedded) and send binary audio chunks.
  4. Read transcript messages (is_final flags partial vs final), send {"type":"stop_recording"} to end, then GET /v2/live/{id} for the full result.

Endpoint

POST https://api.gladia.io/v2/live, then the wss url returned in the response

Authentication

x-gladia-key header on the init POST only; the returned WebSocket URL carries its own session token, so browsers can connect without the API key.

Quick start python

import asyncio, json, os, requests, websockets  # pip install requests websockets>=14

cfg = {"encoding": "wav/pcm", "sample_rate": 16000, "bit_depth": 16, "channels": 1,
       "messages_config": {"receive_partial_transcripts": True}}
r = requests.post("https://api.gladia.io/v2/live", json=cfg,
                  headers={"x-gladia-key": os.environ["GLADIA_API_KEY"]})
r.raise_for_status()
ws_url = r.json()["url"]  # session-scoped, safe to hand to a browser

async def main():
    async with websockets.connect(ws_url) as ws:
        async def send():
            with open("audio_16k_mono.raw", "rb") as f:
                while chunk := f.read(3200):  # 100 ms
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)
            await ws.send(json.dumps({"type": "stop_recording"}))
        asyncio.create_task(send())
        async for msg in ws:
            d = json.loads(msg)
            if d.get("type") == "transcript":
                kind = "FINAL" if d["data"]["is_final"] else "partial"
                print(kind, d["data"]["utterance"]["text"])

asyncio.run(main())

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Audio and transcripts are stored by default

Paid accounts keep audio input and transcripts for 3 weeks by default; free accounts keep everything for 1 year. Zero data retention is Enterprise-only and disables upload and polling endpoints.

Free tier allows only 1 live session

Free accounts get 1 concurrent live session (paid default 30), so any parallel testing needs a paid plan.

Diarization documented for pre-recorded

The speaker diarization page is for pre-recorded audio. Do not assume live speaker labels; use multi-channel input to separate parties on calls.

3-hour session cap

Live sessions are terminated after 3 hours. Long events need a new session and stitching of transcripts.

Format declared up front

encoding, sample_rate, bit_depth and channels must exactly match what you send or the transcript degrades or fails; there is no auto-detection for raw PCM.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Concurrent live sessions: Free 1, Paid 30 (default), Enterprise on demand; HTTP 429 when exhausted.
  • A single live session cannot exceed 3 hours.
  • Paid accounts can also queue up to 300 async jobs (not relevant to live).

Models and products

NameStatusNotes
solaria-1GA (default, only live model)100+ languages, code-switching, async and live.
solaria-3GA (pre-recorded only)EN/FR/DE/ES/IT business audio; not available for live.

Docs and sources

Docs

Sources used

Not fully verified

Exact WebSocket message field names in the snippet (data.is_final, data.utterance.text, stop_recording) are from the v2 API as previously documented and were not re-read line by line. Region and latency info not captured.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?