GA Groq

Groq Whisper (no streaming)

Extremely fast but file-based Whisper transcription. There is no streaming/WebSocket STT endpoint; 'realtime' use means chunking audio yourself and paying a 10-second minimum per request.

Est. per minute$0.00067 - 0.0019
1 high-severity warning

Overview

Best for: Cheap near-real-time transcription of short VAD-segmented utterances where partials are not needed.

At a glance

$/hour$0.04
Free tierYes
Live speakersNo
Turn detectNo
PartialsNo
WebRTCNo
WebSocketNo
gRPCNo
Self-hostNo
Open weightsYes

Not a streaming API: file upload over HTTPS only. $0.04/hr is whisper-large-v3-turbo; large-v3 is $0.111/hr. 10 s minimum billed per request, so short chunks cost more. Open weights refers to the Whisper models; the Groq service is hosted only.

Audio in

flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm files or URL; 25 MB free tier / 100 MB dev tier.

Audio out

n/a

Languages

Multilingual (Whisper).

Latency

Vendor: real-time speed factor 189x (v3) to 216x (turbo) for files; not a streaming latency.

Regions

Groq cloud (not detailed).

Compliance

Not re-verified.

Features

  • segment and word timestamps
  • translation to English (large-v3 only)
  • OpenAI-compatible endpoint

Pricing

WhatPriceUnitNotes
whisper-large-v3-turbo$0.04per hourMinimum billed length 10 s per request.
whisper-large-v3$0.111per hourMinimum billed length 10 s per request.
How the per-minute estimate was worked out

File pricing; chunking live audio into short requests raises effective cost because each request bills at least 10 s.

Free tier: Groq free tier with rate limits (not detailed here).

Source: console.groq.com

Setup

  1. Get a Groq API key.
  2. POST audio to https://api.groq.com/openai/v1/audio/transcriptions (OpenAI SDK compatible).
  3. For pseudo-live use, cut audio on VAD pauses and send each utterance as a file.

Endpoint

https://api.groq.com/openai/v1/audio/transcriptions

Authentication

Authorization: Bearer <GROQ_API_KEY>

Quick start python

import os
from openai import OpenAI  # pip install openai

client = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
with open("utterance.wav", "rb") as f:  # one VAD-segmented utterance, not a live stream
    r = client.audio.transcriptions.create(model="whisper-large-v3-turbo", file=f)
print(r.text)

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Not a streaming API

Groq has only file transcription and translation endpoints. There are no partial results; you must segment audio yourself and accept utterance-level latency.

10-second billing floor

Requests shorter than 10 s are billed as 10 s, so chopping speech into 1-2 s chunks can multiply cost 5-10x.

Whisper hallucinations on silence

Sending silent or noise-only chunks to Whisper models can produce invented text; gate requests with a VAD.

Turbo cannot translate

whisper-large-v3-turbo supports transcription only; translation needs whisper-large-v3.

Plus 12 warnings that apply to all speech-to-text, live APIs. See category warnings.

Limits

  • No streaming endpoint.
  • 10-second minimum billed length per request.
  • Free tier file size 25 MB.

Models and products

NameStatusNotes
whisper-large-v3-turboGA$0.04/hr, transcription only (no translation), 216x real-time speed factor (vendor).
whisper-large-v3GA$0.111/hr, transcription and translation.

Docs and sources

Docs

Sources used

Not fully verified

Rate limits and compliance.

Similar speech-to-text APIs

Spotted a wrong price or a dead link?