Head to head · Speech stt · October 2026 research run

Groq Speech-to-Text vs Soniox Speech-to-Text

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Soniox Speech-to-Text's 58.2 (C), and leads in 6 of 7 scored categories. Both do speech stt.

Which one, for what

Groq Speech-to-Text BB

Good for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.

Ahead on

  • Reliability, 90 against 65
  • Agent ergonomics, 75 against 70
  • Security & auth, 79 against 70
  • Payments & pricing, 40 against 20
  • Maintenance & community, 68 against 30
  • Transparency & trust, 85 against 76

Also in its favour

  • Agent-ready, a grade of BB or better
  • Free to start without a card

Watch for

No streaming or realtime endpoint and no diarisation in the reviewed documentation

Soniox Speech-to-Text C

Good for Price-led multilingual transcription and translation, live or async, and for operators who want no training and no retention.

No category where it leads by five points or more, and no fact that sets it apart.

Watch for

No free credits for new accounts since October 2025

Score by category

CategoryWeight this runGroq Speech-to-TextSoniox Speech-to-TextEdge
Reliability16%209065Groq Speech-to-Text +25
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25860Soniox Speech-to-Text +2
Agent ergonomics13%16.27570Groq Speech-to-Text +5
Security & auth14%17.57970Groq Speech-to-Text +9
Payments & pricing10%12.54020Groq Speech-to-Text +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.86830Groq Speech-to-Text +38
Transparency & trust7%8.88576Groq Speech-to-Text +9
Negative events≤1500
Total71.8 · BB58.2 · C

Facts side by side

FactGroq Speech-to-TextSoniox Speech-to-Text
KindModel APIModel API
VendorGroqSoniox
Hosted endpointhttps://api.groq.com/openai/v1https://api.soniox.com/v1
TransportsHTTPHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
Price for speech sttnot published$0.0017 per minute of audio
x402nono
LicenceProprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging FaceApache-2.0 (Python SDK)
Read-only variant documentednono
llms.txtyesyes
Last release2026-08-262026-08-11
Terms last updated2026-06-222026-06-29
Privacy policy last updated2025-11-122026-06-29
Customer content may train modelsnot found in the textnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waivernot found in the textnot found in the text
Popularity621 stars12 stars, 22k npm/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Groq Speech-to-Text

Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.

Soniox Speech-to-Text

About $0.10 an hour async and $0.12 real time, with diarisation, language ID and translation included. No free credits for new accounts since October 2025.

Before you call either

Groq Speech-to-Text

  1. Send whisper-large-v3-turbo for transcription and whisper-large-v3 for translation to English. The translations endpoint does not accept Turbo.
  2. Pass url instead of file for audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending.
  3. Set response_format to verbose_json before asking for timestamp_granularities[]. Word timestamps add latency, segment timestamps do not.
  4. Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows.
  5. Read retry-after on a 429 and back off. Audio limits count seconds an hour and a day as well as requests.

Soniox Speech-to-Text

  1. Pass audio_url for public files and skip the upload step. Delete uploaded files or they count against the 10 GB quota for 30 days
  2. Use the context field for names and domain terms
  3. Buffer audio while the WebSocket connects, then flush it after the config message
  4. Split anything over 300 minutes. The cap is fixed
  5. Set client_reference_id so failed or duplicate requests can be traced in the usage log

Questions

Which is better for AI agents, Groq Speech-to-Text or Soniox Speech-to-Text?

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Soniox Speech-to-Text's 58.2 (C), and leads in 6 of 7 scored categories.

Do Groq Speech-to-Text and Soniox Speech-to-Text need an API key?

Both need an API key.

Can an agent call Groq Speech-to-Text and Soniox Speech-to-Text without installing anything?

Yes. Groq Speech-to-Text has a hosted endpoint at https://api.groq.com/openai/v1 and Soniox Speech-to-Text at https://api.soniox.com/v1.

Other comparisons with Groq Speech-to-Text or Soniox Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.