Head to head · Speech stt · October 2026 research run

Groq Speech-to-Text vs Mistral Voxtral Transcribe

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation. Both do speech stt.

Which one, for what

Groq Speech-to-Text BB

Good for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.

Ahead on

  • Reliability, 90 against 53
  • Security & auth, 79 against 70

Also in its favour

  • Agent-ready, a grade of BB or better
  • Free to start without a card
  • No incidents deducted, where Mistral Voxtral Transcribe loses 3 points for them

Watch for

No streaming or realtime endpoint and no diarisation in the reviewed documentation

Mistral Voxtral Transcribe B

Good for Suited to low-cost batch transcription with diarisation in the 13 supported languages, and to live captions or voice agents that can run without speaker labels.

Ahead on

  • Schema & documentation, 85 against 58

Watch for

Audio rate limits are named (audio seconds a minute and a month) but the numbers are shown only in the Admin Panel

Score by category

CategoryWeight this runGroq Speech-to-TextMistral Voxtral TranscribeEdge
Reliability16%209053Groq Speech-to-Text +37
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25885Mistral Voxtral Transcribe +27
Agent ergonomics13%16.27577Mistral Voxtral Transcribe +2
Security & auth14%17.57970Groq Speech-to-Text +9
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86866Groq Speech-to-Text +2
Transparency & trust7%8.88582Groq Speech-to-Text +3
Negative events≤150-3
Total71.8 · BB64.1 · B

Facts side by side

FactGroq Speech-to-TextMistral Voxtral Transcribe
KindModel APIModel API
VendorGroqMistral AI
Hosted endpointhttps://api.groq.com/openai/v1https://api.mistral.ai/v1
TransportsHTTPHTTP, websocket
AuthAPI keyAPI key
PricingFreemiumFreemium
x402nono
LicenceProprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging FaceProprietary hosted service under Mistral's commercial terms. Voxtral Mini 4B Realtime weights and the SDKs are Apache-2.0
Read-only variant documentednono
llms.txtyesyes
Last release2026-08-262026-02-04
Terms last updated2026-06-222026-09-25
Privacy policy last updated2025-11-122026-09-03
Customer content may train modelsnot found in the textyes, with an opt-out
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waivernot found in the textnot found in the text
Popularity621 stars773 stars

Verdicts

Groq Speech-to-Text

Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.

Mistral Voxtral Transcribe

Batch transcription costs $0.003 a minute and takes one multipart call with a file, a URL or an uploaded file ID. Rate limit numbers are shown only in the console, and timestamps cannot be combined with a set language, nor diarisation with the realtime model.

Before you call either

Groq Speech-to-Text

  1. Send whisper-large-v3-turbo for transcription and whisper-large-v3 for translation to English. The translations endpoint does not accept Turbo.
  2. Pass url instead of file for audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending.
  3. Set response_format to verbose_json before asking for timestamp_granularities[]. Word timestamps add latency, segment timestamps do not.
  4. Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows.
  5. Read retry-after on a 429 and back off. Audio limits count seconds an hour and a day as well as requests.

Mistral Voxtral Transcribe

  1. Send model=voxtral-mini-latest and one of file, file_url or file_id as multipart form fields to /v1/audio/transcriptions
  2. Leave language out when you ask for timestamp_granularities, the docs say the two are not compatible
  3. Use the batch endpoint for diarisation. voxtral-mini-transcribe-realtime-2602 does not accept diarize
  4. Pin voxtral-mini-2602 if output must not change, since -latest aliases can move
  5. Mint browser tokens with POST /v1/client/sessions close to connection time and pass them in Sec-WebSocket-Protocol, never the API key

Questions

Which is better for AI agents, Groq Speech-to-Text or Mistral Voxtral Transcribe?

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Mistral Voxtral Transcribe's 64.1 (B), and leads in 4 of 7 scored categories. Mistral Voxtral Transcribe leads on schema & documentation.

Do Groq Speech-to-Text and Mistral Voxtral Transcribe need an API key?

Both need an API key.

Can an agent call Groq Speech-to-Text and Mistral Voxtral Transcribe without installing anything?

Yes. Groq Speech-to-Text has a hosted endpoint at https://api.groq.com/openai/v1 and Mistral Voxtral Transcribe at https://api.mistral.ai/v1.

Other comparisons with Groq Speech-to-Text or Mistral Voxtral Transcribe

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.