Groq Whisper (no streaming)
Extremely fast but file-based Whisper transcription. There is no streaming/WebSocket STT endpoint; 'realtime' use means chunking audio yourself and paying a 10-second minimum per request.
Overview
Best for: Cheap near-real-time transcription of short VAD-segmented utterances where partials are not needed.
At a glance
Not a streaming API: file upload over HTTPS only. $0.04/hr is whisper-large-v3-turbo; large-v3 is $0.111/hr. 10 s minimum billed per request, so short chunks cost more. Open weights refers to the Whisper models; the Groq service is hosted only.
flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm files or URL; 25 MB free tier / 100 MB dev tier.
n/a
Multilingual (Whisper).
Vendor: real-time speed factor 189x (v3) to 216x (turbo) for files; not a streaming latency.
Groq cloud (not detailed).
Not re-verified.
Features
- segment and word timestamps
- translation to English (large-v3 only)
- OpenAI-compatible endpoint
Pricing
| What | Price | Unit |
|---|---|---|
| whisper-large-v3-turbo | $0.04 | per hour |
| whisper-large-v3 | $0.111 | per hour |
File pricing; chunking live audio into short requests raises effective cost because each request bills at least 10 s.
Free tier: Groq free tier with rate limits (not detailed here).
Source: console.groq.com
Setup
- Get a Groq API key.
- POST audio to https://api.groq.com/openai/v1/audio/transcriptions (OpenAI SDK compatible).
- For pseudo-live use, cut audio on VAD pauses and send each utterance as a file.
Endpoint
https://api.groq.com/openai/v1/audio/transcriptions
Authentication
Authorization: Bearer <GROQ_API_KEY>
Quick start python
import os
from openai import OpenAI # pip install openai
client = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
with open("utterance.wav", "rb") as f: # one VAD-segmented utterance, not a live stream
r = client.audio.transcriptions.create(model="whisper-large-v3-turbo", file=f)
print(r.text)
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Not a streaming API
Groq has only file transcription and translation endpoints. There are no partial results; you must segment audio yourself and accept utterance-level latency.
10-second billing floor
Requests shorter than 10 s are billed as 10 s, so chopping speech into 1-2 s chunks can multiply cost 5-10x.
Whisper hallucinations on silence
Sending silent or noise-only chunks to Whisper models can produce invented text; gate requests with a VAD.
Turbo cannot translate
whisper-large-v3-turbo supports transcription only; translation needs whisper-large-v3.
Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.
Limits
- No streaming endpoint.
- 10-second minimum billed length per request.
- Free tier file size 25 MB.
Models and products
| Name | Status |
|---|---|
| whisper-large-v3-turbo | GA |
| whisper-large-v3 | GA |
Docs and sources
Docs
Sources used
Rate limits and compliance.