Head to head · Speech stt · October 2026 research run
Groq Speech-to-Text vs Mistral Voxtral Transcribe
Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation. Both do speech stt.
Which one, for what
Good for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.
Ahead on
- Reliability, 90 against 53
- Security & auth, 79 against 70
Also in its favour
- Agent-ready, a grade of BB or better
- Free to start without a card
- No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them
Watch for
No streaming or realtime endpoint and no diarisation in the reviewed documentation
Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.
Ahead on
- Schema & documentation, 85 against 58
Watch for
Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel
Score by category
| Category | Weight this run | Groq Speech-to-Text | Mistral Voxtral Transcribe | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 90 | 53 | Groq Speech-to-Text +37 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 58 | 85 | Mistral Voxtral Transcribe +27 |
| Agent ergonomics | 13%16.2 | 75 | 77 | Mistral Voxtral Transcribe +2 |
| Security & auth | 14%17.5 | 79 | 70 | Groq Speech-to-Text +9 |
| Payments & pricing | 10%12.5 | 40 | 40 | even |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 68 | 66 | Groq Speech-to-Text +2 |
| Transparency & trust | 7%8.8 | 85 | 82 | Groq Speech-to-Text +3 |
| Negative events | ≤15 | 0 | -3 | |
| Total | 71.8 · BB | 64.1 · B |
Facts side by side
| Fact | Groq Speech-to-Text | Mistral Voxtral Transcribe |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Groq | Mistral AI |
| Hosted endpoint | https://api.groq.com/openai/v1 | https://api.mistral.ai/v1 |
| Transports | HTTP | HTTP, websocket |
| Auth | API key | API key |
| Pricing | Freemium | Freemium |
| x402 | no | no |
| Licence | Proprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging Face | Proprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0 |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-08-26 | 2026-02-04 |
| Terms last updated | 2026-06-22 | 2026-09-25 |
| Privacy policy last updated | 2025-11-12 | 2026-09-03 |
| Customer content may train models | not found in the text | yes, with an opt-out |
| Terms restrict automated access | not found in the text | not found in the text |
| Terms restrict benchmarking | yes | yes |
| Terms or service can change without notice | not found in the text | yes |
| Arbitration or class-action waiver | not found in the text | not found in the text |
| Popularity | 621 stars | 773 stars |
Verdicts
Groq Speech-to-Text
Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.
Mistral Voxtral Transcribe
Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.
Before you call either
Groq Speech-to-Text
- Send
whisper-large-v3-turbofor transcription andwhisper-large-v3for translation to English. The translations endpoint does not accept Turbo. - Pass
urlinstead offilefor audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending. - Set
response_formattoverbose_jsonbefore asking fortimestamp_granularities[]. Word timestamps add latency, segment timestamps do not. - Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows.
- Read
retry-afteron a 429 and back off. Audio limits count seconds an hour and a day as well as requests.
Mistral Voxtral Transcribe
- Send
model=voxtral-mini-latestand one offile,file_urlorfile_idas multipart form fields to/v1/audio/transcriptions - Leave
languageout when you ask fortimestamp_granularities, the docs say the two are not compatible - Use the batch endpoint for diarisation.
voxtral-mini-transcribe-realtime-2602does not acceptdiarize - Pin
voxtral-mini-2602if output must not change, since-latestaliases can move - Mint browser tokens with
POST /v1/client/sessionsclose to connection time and pass them inSec-WebSocket-Protocol, never the API key
Questions
Which is better for AI agents, Groq Speech-to-Text or Mistral Voxtral Transcribe?
Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation.
Do Groq Speech-to-Text and Mistral Voxtral Transcribe need an API key?
Both need an API key.
Can an agent call Groq Speech-to-Text and Mistral Voxtral Transcribe without installing anything?
Yes. Groq Speech-to-Text has a hosted endpoint at https://api.groq.com/openai/v1 and Mistral Voxtral Transcribe at https://api.mistral.ai/v1.
Other comparisons with Groq Speech-to-Text or Mistral Voxtral Transcribe
- Amazon Transcribe vs Groq Speech-to-Text
- Amazon Transcribe vs Mistral Voxtral Transcribe
- AssemblyAI Speech-to-Text (Universal) vs Groq Speech-to-Text
- AssemblyAI Speech-to-Text (Universal) vs Mistral Voxtral Transcribe
- Azure AI Speech speech-to-text vs Groq Speech-to-Text
- Azure AI Speech speech-to-text vs Mistral Voxtral Transcribe
- Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe
- ElevenLabs Scribe Speech to Text API vs Groq Speech-to-Text
- ElevenLabs Scribe Speech to Text API vs Mistral Voxtral Transcribe
- Gladia Speech-to-Text API + MCP vs Groq Speech-to-Text
- Gladia Speech-to-Text API + MCP vs Mistral Voxtral Transcribe
- Google Cloud Speech-to-Text vs Groq Speech-to-Text
- Google Cloud Speech-to-Text vs Mistral Voxtral Transcribe
- Groq Speech-to-Text vs Rev AI Speech-to-Text API
- Groq Speech-to-Text vs Soniox Speech-to-Text
- Groq Speech-to-Text vs Speechmatics Speech-to-Text
- Mistral Voxtral Transcribe vs Rev AI Speech-to-Text API
- Mistral Voxtral Transcribe vs Soniox Speech-to-Text
- Mistral Voxtral Transcribe vs Speechmatics Speech-to-Text
Machine-readable
- This page as Markdown
/compare/groq-speech-to-text-vs-mistral-voxtral-transcribe.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/groq-speech-to-text.json·/api/v1/tools/mistral-voxtral-transcribe.json - From a terminal
anchor compare groq-speech-to-text mistral-voxtral-transcribe(the CLI) - Over MCP
compare_tools {"a": "groq-speech-to-text", "b": "mistral-voxtral-transcribe"}at/mcp, no key