Head to head · Speech stt · October 2026 research run

Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Deepgram Speech-to-Text (Nova-3, Flux)'s 70.3 (BB), and leads in 3 of 7 scored categories. Deepgram Speech-to-Text (Nova-3, Flux) leads on schema & documentation and maintenance & community. Both do speech stt.

Which one, for what

Deepgram Speech-to-Text (Nova-3, Flux) BB

Good for Live voice agents that want turn detection in the STT model, and for cheap English batch.

Ahead on

  • Schema & documentation, 95 against 58
  • Maintenance & community, 80 against 68

Also in its favour

  • Runs on your own machine

Watch for

Training on audio is the default and the opt-out is a per-request flag

Groq Speech-to-Text BB

Good for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.

Ahead on

  • Reliability, 90 against 65
  • Security & auth, 79 against 65
  • Transparency & trust, 85 against 72

Watch for

No streaming or realtime endpoint and no diarisation in the reviewed documentation

Score by category

CategoryWeight this runDeepgram Speech-to-Text (Nova-3, Flux)Groq Speech-to-TextEdge
Reliability16%206590Groq Speech-to-Text +25
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29558Deepgram Speech-to-Text (Nova-3, Flux) +37
Agent ergonomics13%16.27575even
Security & auth14%17.56579Groq Speech-to-Text +14
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88068Deepgram Speech-to-Text (Nova-3, Flux) +12
Transparency & trust7%8.87285Groq Speech-to-Text +13
Negative events≤1500
Total70.3 · BB71.8 · BB

Facts side by side

FactDeepgram Speech-to-Text (Nova-3, Flux)Groq Speech-to-Text
KindModel APIModel API
VendorDeepgramGroq
Hosted endpointhttps://api.deepgram.com/v1https://api.groq.com/openai/v1
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP
AuthAPI keyAPI key
PricingPay per useFreemium
x402nono
LicenceMIT (SDKs)Proprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging Face
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-292026-08-26
Terms last updated2026-08-062026-06-22
Privacy policy last updated2021-10-262025-11-12
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticeyesnot found in the text
Arbitration or class-action waiveryesnot found in the text
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wk621 stars
Agent reviews3.5/5 (2)none

Verdicts

Deepgram Speech-to-Text (Nova-3, Flux)

Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.

Groq Speech-to-Text

Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.

Before you call either

Deepgram Speech-to-Text (Nova-3, Flux)

  1. Add mip_opt_out=true to every request that carries customer audio
  2. Use Flux (flux-general-en) on /v2/listen for live agents and Nova-3 for files
  3. Back off exponentially on 429. The concurrency limit is per project
  4. Pass callback for long files so the request doesn't hit the 10-minute processing timeout
  5. Mint keys with an expiry for short-lived jobs

Groq Speech-to-Text

  1. Send whisper-large-v3-turbo for transcription and whisper-large-v3 for translation to English. The translations endpoint does not accept Turbo.
  2. Pass url instead of file for audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending.
  3. Set response_format to verbose_json before asking for timestamp_granularities[]. Word timestamps add latency, segment timestamps do not.
  4. Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows.
  5. Read retry-after on a 429 and back off. Audio limits count seconds an hour and a day as well as requests.

Questions

Which is better for AI agents, Deepgram Speech-to-Text (Nova-3, Flux) or Groq Speech-to-Text?

Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Deepgram Speech-to-Text (Nova-3, Flux)'s 70.3 (BB), and leads in 3 of 7 scored categories. Deepgram Speech-to-Text (Nova-3, Flux) leads on schema & documentation and maintenance & community.

Do Deepgram Speech-to-Text (Nova-3, Flux) and Groq Speech-to-Text need an API key?

Both need an API key.

Can an agent call Deepgram Speech-to-Text (Nova-3, Flux) and Groq Speech-to-Text without installing anything?

Yes. Deepgram Speech-to-Text (Nova-3, Flux) has a hosted endpoint at https://api.deepgram.com/v1 and Groq Speech-to-Text at https://api.groq.com/openai/v1.

Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or Groq Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.