GA Moonshine AI (Useful Sensors)

Moonshine Voice (Moonshine v2 streaming)

MIT-licensed on-device STT library with streaming models (Moonshine v2, 2026) that do work while the user is still talking. Runs on CPU, phones, browsers (WASM) and Raspberry Pi.

Est. per minute$0

Overview

Best for: Private, offline, low-latency captions and voice commands on phones, browsers and single-board computers.

At a glance

TypeSpeech-to-text
CPU okYes
LicenceMIT
CommercialYes
StreamingYes

On-device library; model sizes range from tiny to large. English plus other languages (list not captured). Legacy non-English non-streaming models are non-commercial.

Audio in

Microphone or PCM via the library.

Audio out

n/a

Languages

English plus additional languages (list on docs site; not captured).

Latency

Paper: bounded time-to-first-token independent of utterance length; third-party claim of sub-200 ms on edge devices is unverified.

Regions

On-device.

Compliance

On-device processing.

Hardware

CPU-only friendly: laptops, phones, Raspberry Pi, browser WASM.

Licence

MIT (code and streaming models); legacy non-English non-streaming models non-commercial.

Features

  • streaming partials
  • on-device
  • intent recognition and TTS in the same library
  • cross-platform

Pricing

WhatPriceUnitNotes
Library and models$0MIT, except legacy non-English non-streaming models.
How the per-minute estimate was worked out

Runs on the user's device.

Free tier: Free and open source.

Source: github.com

Setup

  1. pip install moonshine-voice
  2. moonshine-voice mic --language en (live mic test)
  3. See moonshine-voice.readthedocs.io for Python/JS/mobile APIs.

Endpoint

n/a (local)

Authentication

none

Quick start bash

pip install moonshine-voice
moonshine-voice mic --language en   # live transcription from the default microphone

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Licence exception

Everything is MIT except the legacy non-streaming models for non-English languages, which stay under a non-commercial Moonshine licence. Use the streaming models for commercial products.

Benchmarks disagree

The arXiv abstract and the newer PDF on moonshine.ai give different accuracy/speed claims (6x vs 4x model size parity). Test on your own audio and hardware.

Small models trade accuracy

Tiny and micro models suit commands and captions on weak hardware; long-form dictation needs the larger models.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • Accuracy depends on chosen model size and device CPU.

Models and products

NameStatusNotes
Moonshine v2 streaming modelsReleased 2026Sliding-window attention encoder for bounded latency (arXiv 2602.12241). Sizes range from tiny (~1 MB micro) up to models the README says beat Whisper Large V3.
Legacy non-streaming non-English modelsLegacyUnder a non-commercial Moonshine licence (exception to MIT).

Docs and sources

Docs

Sources used

Not fully verified

Supported language list and per-device latency numbers.

Similar open models APIs

Spotted a wrong price or a dead link?