GA Liquid AI

LFM2.5-Audio-1.5B

A tiny (1.5B) end-to-end speech and text model that does interleaved speech-to-speech chat and runs on CPU via GGUF. Ideal for on-device or edge voice, as long as your company is under the $10M revenue licence threshold.

Est. per minuten/a
1 high-severity warning

Overview

Best for: On-device or edge voice assistants for small companies, offline kiosks, prototypes.

At a glance

TypeVoice-to-voice
Params B1.5
CPU okYes
LicenceCustom
CommercialNo
Full duplexNo
Languages1

LFM Open License v1.0: commercial use only for entities under $10M annual revenue. English (separate Japanese variant). Runs on CPU via GGUF.

Audio in

Speech

Audio out

Speech (interleaved with text)

Languages

English (Japanese variant available).

Voices

Not documented.

Latency

Designed for low-latency real-time conversation; no figure.

Regions

Self-hosted / on-device.

Compliance

Self-hosted.

Hardware

Runs on CPU via GGUF; a small GPU speeds it up (bfloat16, flash-attn optional).

Licence

LFM Open License v1.0: commercial use only for entities under $10M annual revenue; above that, use is not licensed.

Features

  • interleaved speech-to-speech chat
  • sequential ASR/TTS modes
  • CPU inference via GGUF

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. Check the licence threshold first.
  2. pip install liquid-audio and launch the demo, or use the GGUF build with llama.cpp on CPU.

Endpoint

Local Gradio demo on port 7860

Authentication

n/a

Quick start python

pip install liquid-audio
pip install "liquid-audio[demo]"
liquid-audio-demo   # serves on http://localhost:7860
# CPU: use LiquidAI/LFM2.5-Audio-1.5B-GGUF with llama.cpp

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Revenue cap in the licence

Commercial use is licensed only if you (or your legal entity) have under $10,000,000 annual revenue. Larger companies need a separate deal with Liquid AI.

Tiny model limits

1.5B parameters means weak knowledge and reasoning; pair with retrieval or keep tasks narrow.

English only

Main model is English; a separate JP model exists.

No full-duplex server

Interleaved generation is turn-based; barge-in handling is up to you.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Small model: limited knowledge/reasoning
  • English only (main model)

Models and products

NameStatusNotes
LiquidAI/LFM2.5-Audio-1.5B (+ GGUF, ONNX)GAAbout 1.2B LM params, 1.5B total.
LiquidAI/LFM2.5-Audio-1.5B-JPGAJapanese variant (May 2026).
LiquidAI/LFM2-Audio-1.5BGAPrevious version.

Docs and sources

Docs

Sources used

Not fully verified

Latency numbers; exact llama.cpp invocation.

Similar open models APIs

Spotted a wrong price or a dead link?