Smallest.ai Waves Lightning v3.1
44.1 kHz TTS with 217 voices, instant cloning, and HTTP/SSE/WebSocket transports. Default account concurrency is just 1 active TTS request.
Overview
Best for: Indian-language and English agents wanting cloning and 44.1 kHz audio, once concurrency is raised.
At a glance
200 ms TTFB at 40 concurrent in-region (vendor); Vapi measured 420 ms. Official per-character price not published; third parties say ~$14.50/1M. 12 trained languages (20 codes accepted). Default concurrency 1 per account. Hosts in India and US only.
Text (pronunciation dictionaries on WebSocket)
pcm default; 8000/16000/24000/44100 Hz
12 with trained voices; 20 accepted codes plus auto
217 catalog voices; instant cloning from 5-15 s
Vendor: 200 ms TTFB at 40 concurrent requests in-region. Vapi measured 420 ms median including network.
India (Mumbai, api.india.smallest.ai) and USA (Oregon, api.us.smallest.ai), auto-routed
Zero data retention on enterprise plans.
Features
- input streaming with context_id continuations
- word timestamps (WebSocket only)
- voice cloning
- pronunciation dictionaries
- zero data retention (enterprise)
- region-pinned hosts (India, US)
Pricing
| What | Price | Unit |
|---|---|---|
| Lightning v3.1 | Not confirmed |
Official per-character rate not retrievable; third-party figures imply ~$0.013-0.023/min
Free tier: $10 free credits (platform)
Source: smallest.ai
Setup
- Create an API key in the Waves dashboard.
- Connect to wss://api.smallest.ai/waves/v1/tts/live with Authorization: Bearer.
- Send JSON requests with text, voice_id, sample_rate, output_format, optionally context_id for continuations.
- Decode base64 data.audio messages.
Endpoint
wss://api.smallest.ai/waves/v1/tts/live
Authentication
Authorization: Bearer <SMALLEST_API_KEY>
Quick start javascript
// npm i ws - shape follows the Lightning v3.1 model card example
import WebSocket from "ws";
import fs from "fs";
const ws = new WebSocket("wss://api.smallest.ai/waves/v1/tts/live", {
headers: { Authorization: `Bearer ${process.env.SMALLEST_API_KEY}` },
});
const out = fs.createWriteStream("out_24k.pcm");
ws.on("open", () => {
ws.send(JSON.stringify({ text: "Hello from Lightning.", voice_id: "YOUR_VOICE_ID",
sample_rate: 24000, output_format: "pcm" }));
});
ws.on("message", (raw) => {
const msg = JSON.parse(raw.toString());
if (msg.data?.audio) out.write(Buffer.from(msg.data.audio, "base64"));
if (msg.status === "complete") ws.close();
});
ws.on("close", () => out.end());
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Concurrency of 1 by default
Only one TTS request can be processing per account at a time; a second returns 429. Vendor suggests ~4 parallel conversations at best. Get an enterprise limit before production.
Pricing is unclear
The pricing page does not list a Lightning per-character rate; third-party numbers disagree. Get a quote.
context_id cannot combine with flush
Continuations via context_id cannot be combined with flush or max_buffer_flush_ms.
8 of 20 language codes reuse other voices
Only 12 languages have trained voices; the other 8 route through English or Hindi voices.
Plus 12 warnings that apply to all text-to-speech, streaming APIs. See category warnings.
Limits
- Default: 1 concurrent TTS request per account across all endpoints
- Up to 5 WebSocket connections, still 1 active request
- Pay-as-you-go platform concurrency listed as 20 for agents
Models and products
| Name | Status |
|---|---|
| lightning-v3.1 | GA |
| lightning-v3.1-pro | GA |
| lightning-v2 / lightning-large | Deprecated |
Docs and sources
Docs
Sources used
- docs.smallest.ai/models/model-cards/text-to-speech/lightning-v-3-1.md
- docs.smallest.ai/api-reference/concurrency-and-limits.md
- smallest.ai/pricing
- humannessindex.vapi.ai/models/smallestai-lightning-v31
Per-character price; exact response message shape (status field name) in the snippet.