GA Picovoice

Picovoice Cheetah (on-device streaming STT)

Commercial on-device streaming STT SDK running locally on desktop, mobile, web and Raspberry Pi. Leopard is the on-device batch counterpart. Requires a Picovoice AccessKey.

Est. per minute$0.02
1 high-severity warning

Overview

Best for: Offline or privacy-critical apps (kiosks, embedded, mobile) that need streaming STT without sending audio to a cloud.

At a glance

$/hour$1.2
Free tierYes
Turn detectYes
KeytermsYes
PartialsYes
Languages8
WebRTCNo
WebSocketNo
gRPCNo
Self-hostYes
Open weightsNo

On-device SDK, audio never leaves the device. Free tier is non-commercial only. Price is the Foundation Plan ($6,000/yr for 25K min/month) fully used; unused minutes raise the rate. 8 languages, more on Enterprise.

Audio in

16-bit PCM frames of Cheetah's frame_length (16 kHz mono in Picovoice SDKs; not restated on the doc page read).

Audio out

n/a

Languages

English, French, German, Italian, Japanese, Korean, Spanish, Portuguese; more for Enterprise.

Latency

Runs locally; no network round trip. No numeric claim captured.

Regions

On-device; audio never leaves the device.

Compliance

Audio processed on device (privacy by design).

Hardware

Runs on CPU: Linux x86_64, macOS (x86_64, arm64), Windows (x86_64, arm64), Android, iOS, Web (WASM), Raspberry Pi 3/4/5.

Licence

Proprietary SDK; AccessKey required.

Features

  • partial transcripts
  • endpoint detection (is_endpoint flag in SDK)
  • automatic punctuation (option)
  • custom vocabulary with optional IPA
  • keyword boosting

Pricing

WhatPriceUnitNotes
Free tier$0Personal / non-commercial use; FAQ mentions up to 5 hours/month across listed engines. Third-party claim that the free tier ends 2026-06-30 is unconfirmed.
Foundation Plan (startups)$6,000per yearStartups under 5 years old and 20 employees or fewer; includes 25K Cheetah minutes/month plus other engines.
Enterprisecustom
How the per-minute estimate was worked out

Foundation Plan if fully used: $6,000 / (25,000 min x 12) = $0.02/min. Unused minutes raise the effective rate.

Free tier: Non-commercial free tier; commercial free trial grants Foundation Plan rights.

Source: picovoice.ai

Setup

  1. Sign up at Picovoice Console and copy your AccessKey.
  2. pip install pvcheetah pvrecorder (SDKs also for Android, iOS, Web, C, .NET, etc.).
  3. Feed frames of cheetah.frame_length samples; print partials; call flush() at endpoints.

Endpoint

n/a (local library)

Authentication

AccessKey passed to create(); validated against your account.

Quick start python

import os
import pvcheetah, pvrecorder  # pip install pvcheetah pvrecorder

cheetah = pvcheetah.create(access_key=os.environ["PV_ACCESS_KEY"],
                           enable_automatic_punctuation=True)
rec = pvrecorder.PvRecorder(frame_length=cheetah.frame_length)
rec.start()
try:
    while True:
        partial, is_endpoint = cheetah.process(rec.read())
        print(partial, end="", flush=True)
        if is_endpoint:
            print(cheetah.flush())
except KeyboardInterrupt:
    pass
finally:
    rec.stop(); rec.delete(); cheetah.delete()

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Free tier is non-commercial

Picovoice describes the free tier for personal non-commercial projects; shipping a product needs a paid plan (Foundation $6,000/yr for eligible startups, otherwise Enterprise).

AccessKey still meters usage

Even though inference is local, the SDK requires an AccessKey that is checked against account limits; plan for that dependency and key protection in apps.

Limited languages

Eight languages out of the box; others only through Enterprise.

Accuracy below cloud leaders

On-device models trade accuracy for privacy and cost; benchmark against your audio before choosing.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • Usage metered against account limits via AccessKey.
  • Commercial use requires a paid plan.

Models and products

NameStatusNotes
CheetahGAStreaming, on-device, custom vocabulary and keyword boosting via Picovoice Console.
LeopardGAOn-device batch (file) transcription; not streaming.

Docs and sources

Docs

Sources used

Not fully verified

Current free-tier status in 2026 (pricing page did not render), exact SDK signature details (create(enable_automatic_punctuation), process() returning (partial, is_endpoint)) are from prior SDK versions.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?