# Groq Speech-to-Text (slim) > Groq's hosted speech-to-text API. It runs OpenAI's Whisper Large v3 and Whisper Large v3 Turbo on OpenAI-compatible transcription and translation endpoints, for uploaded files or audio URLs, with a half-price batch mode. - Full: https://www.anchorterminal.com/tools/groq-speech-to-text.md (~7,400 tokens) · this version ~1,480 tokens · JSON https://www.anchorterminal.com/tools/groq-speech-to-text.json · canonical https://www.anchorterminal.com/tools/groq-speech-to-text - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **BB · 71.8/100 · rank #113 of 842 · #3 in Speech-to-text · agent-ready · confidence medium** Assessment: Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan. ## Facts - Kind: Model API · vendor: Groq · category: Speech-to-text · legal entity: Groq LLC · provenance 98/100 - Endpoint: `https://api.groq.com/openai/v1` (HTTP) - Auth: API key · pricing: Freemium · x402: no · licence: Proprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging Face - Probe metrics: not measured yet (probes haven't run) - Models: `whisper-large-v3-turbo` (transcription) and `whisper-large-v3` (transcription and translation to English) - Languages: Vendor says 99+ languages, with an optional ISO-639-1 `language` hint - Max audio: 25 MB a file on the free plan, 100 MB on the Developer plan. Longer audio is chunked by the caller - Input: Multipart `file` or a `url`, in flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav or webm. Only the first audio track is read - Output: `json`, `text` or `verbose_json` with segment or word timestamps. No `srt` or `vtt` - Streaming: None found. One request returns the whole transcript - Diarisation: None found - Rate limits: Free plan 20 requests a minute, 2,000 a day, 7,200 audio seconds an hour and 28,800 a day. Developer plan 300 requests a minute and 200,000 audio seconds an hour on v3, 400 and 400,000 on Turbo - Batch: 50% off, `url` input only, 24-hour to 7-day window - Minimum charge: 10 seconds of audio a request - Free tier: No card, at the free plan's rate limits - Trains on API data: No, barred by the services agreement unless the customer grants permission - Data retention: None by default, up to 30 days when Groq logs for reliability or abuse. Zero retention is a setting in Data Controls - Data location: Google Cloud storage in the US - Prices: Whisper Large v3 Turbo ($0.04 an audio hour) $0.0007 per minute of audio; Whisper Large v3 ($0.111 an audio hour) $0.0019 per minute of audio - Scores: Reliability 90, Performance pending, Schema & documentation 58, Agent ergonomics 75, Security & auth 79, Payments & pricing 40, Task success pending, Maintenance & community 68, Transparency & trust 85 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading. · Agent ergonomics, API reading, as for the other speech-to-text listings. · Security & auth, Model reading, scored like the GroqCloud listing. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Model reading, as for the GroqCloud listing. · Transparency & trust, The hosted service is closed under the Groq Services Agreement. - Sources: 26, open questions: 12, both in the full twin - Capabilities: speech.stt, speech.batch, speech.languages - JSON: https://www.anchorterminal.com/api/v1/tools/groq-speech-to-text.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/groq-speech-to-text.svg` or a link to https://www.anchorterminal.com/tools/groq-speech-to-text from a page on groq.com or one of its subdomains, or the README of github.com/groq/groq-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send `whisper-large-v3-turbo` for transcription and `whisper-large-v3` for translation to English. The translations endpoint does not accept Turbo. 2. Pass `url` instead of `file` for audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending. 3. Set `response_format` to `verbose_json` before asking for `timestamp_granularities[]`. Word timestamps add latency, segment timestamps do not. 4. Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows. 5. Read `retry-after` on a 429 and back off. Audio limits count seconds an hour and a day as well as requests. ## Connect ```bash pip install groq # or: npm install --save groq-sdk ``` ```bash curl https://api.groq.com/openai/v1/audio/transcriptions \ -H "Authorization: Bearer $GROQ_API_KEY" \ -H "Content-Type: multipart/form-data" \ -F file="@./sample_audio.m4a" \ -F model="whisper-large-v3" ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Amazon Transcribe | BB | 73.4 | speech.stt, speech.batch, speech.languages | https://www.anchorterminal.com/tools/amazon-transcribe.min.md | | Azure AI Speech speech-to-text | BB | 73 | speech.stt, speech.batch, speech.languages | https://www.anchorterminal.com/tools/azure-speech-to-text.min.md | | Deepgram Speech-to-Text (Nova-3, Flux) | BB | 70.3 | speech.stt, speech.batch, speech.languages | https://www.anchorterminal.com/tools/deepgram-stt.min.md | | Google Cloud Speech-to-Text | BB | 70.2 | speech.stt, speech.batch, speech.languages | https://www.anchorterminal.com/tools/google-speech-to-text.min.md | | Gladia Speech-to-Text API + MCP | B | 69.5 | speech.stt, speech.batch, speech.languages | https://www.anchorterminal.com/tools/gladia-stt.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)