Head to head · Speech stt · October 2026 research run

Groq Speech-to-Text vs Speechmatics Speech-to-Text

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Speechmatics Speech-to-Text's 67.1 (B), and leads in 3 of 7 scored categories. Speechmatics Speech-to-Text leads on schema & documentation, agent ergonomics and maintenance & community. Both do speech stt.

Which one, for what

Groq Speech-to-Text BB

Good for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.

Ahead on

  • Reliability, 90 against 70
  • Security & auth, 79 against 65
  • Transparency & trust, 85 against 75

Also in its favour

  • Agent-ready, a grade of BB or better

Watch for

No streaming or realtime endpoint and no diarisation in the reviewed documentation

Speechmatics Speech-to-Text B

Good for Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns.

Ahead on

  • Schema & documentation, 65 against 58
  • Agent ergonomics, 80 against 75
  • Maintenance & community, 75 against 68

Watch for

Enhanced costs $0.40 to $0.43 an hour, above most rivals

Score by category

CategoryWeight this runGroq Speech-to-TextSpeechmatics Speech-to-TextEdge
Reliability16%209070Groq Speech-to-Text +20
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25865Speechmatics Speech-to-Text +7
Agent ergonomics13%16.27580Speechmatics Speech-to-Text +5
Security & auth14%17.57965Groq Speech-to-Text +14
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86875Speechmatics Speech-to-Text +7
Transparency & trust7%8.88575Groq Speech-to-Text +10
Negative events≤1500
Total71.8 · BB67.1 · B

Facts side by side

FactGroq Speech-to-TextSpeechmatics Speech-to-Text
KindModel APIModel API
VendorGroqSpeechmatics
Hosted endpointhttps://api.groq.com/openai/v1https://eu1.asr.api.speechmatics.com/v2
TransportsHTTPHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
Price for speech sttnot published$0.0027 per minute of audio
x402nono
LicenceProprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging FaceMIT (SDKs)
Read-only variant documentednono
llms.txtyesyes
Last release2026-08-262026-09-22
Terms last updated2026-06-22no date given
Privacy policy last updated2025-11-122026-05-27
Customer content may train modelsnot found in the textnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waivernot found in the textnot found in the text
Popularity621 stars20 stars, 58k npm/wk, 48k PyPI/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Groq Speech-to-Text

Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.

Speechmatics Speech-to-Text

Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour.

Before you call either

Groq Speech-to-Text

  1. Send whisper-large-v3-turbo for transcription and whisper-large-v3 for translation to English. The translations endpoint does not accept Turbo.
  2. Pass url instead of file for audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending.
  3. Set response_format to verbose_json before asking for timestamp_granularities[]. Word timestamps add latency, segment timestamps do not.
  4. Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows.
  5. Read retry-after on a 429 and back off. Audio limits count seconds an hour and a day as well as requests.

Speechmatics Speech-to-Text

  1. Set "model": "enhanced" explicitly. The default is standard
  2. Use notifications instead of polling. Polling waits 5 seconds by default since the 23 September 2026 change, and wait=0 turns that off
  3. Fetch batch transcripts within 7 days. After that the API returns 404 expired
  4. Pass a fetch_data URL for files over 1 GB
  5. Use /v2/agent with linden-1 for live agents instead of the plain realtime path

Questions

Which is better for AI agents, Groq Speech-to-Text or Speechmatics Speech-to-Text?

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Speechmatics Speech-to-Text's 67.1 (B), and leads in 3 of 7 scored categories. Speechmatics Speech-to-Text leads on schema & documentation, agent ergonomics and maintenance & community.

Do Groq Speech-to-Text and Speechmatics Speech-to-Text need an API key?

Both need an API key.

Can an agent call Groq Speech-to-Text and Speechmatics Speech-to-Text without installing anything?

Yes. Groq Speech-to-Text has a hosted endpoint at https://api.groq.com/openai/v1 and Speechmatics Speech-to-Text at https://eu1.asr.api.speechmatics.com/v2.

Other comparisons with Groq Speech-to-Text or Speechmatics Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.