Head to head · Speech stt · October 2026 research run

Deepgram Speech-to-Text (Nova-3, Flux) vs Speechmatics Speech-to-Text

Deepgram Speech-to-Text (Nova-3, Flux) has a score of 70.6 (BB) against Speechmatics Speech-to-Text's 67.3 (B). Both do speech stt. The largest gap is schema & documentation, 30 points.

Which one, for what

Pick Deepgram Speech-to-Text (Nova-3, Flux) for

  • schema & documentation (+30)
  • maintenance & community (+5)

Pick Speechmatics Speech-to-Text for

  • reliability (+5)
  • agent ergonomics (+5)

Score by category

CategoryWeight this runDeepgram Speech-to-Text (Nova-3, Flux)Speechmatics Speech-to-TextEdge
Reliability16%206570Speechmatics Speech-to-Text +5
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29565Deepgram Speech-to-Text (Nova-3, Flux) +30
Agent ergonomics13%16.27580Speechmatics Speech-to-Text +5
Security & auth14%17.56565even
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88075Deepgram Speech-to-Text (Nova-3, Flux) +5
Transparency & trust7%8.87578Speechmatics Speech-to-Text +3
Negative events≤1500
Total70.6 · BB67.3 · B

Facts side by side

FactDeepgram Speech-to-Text (Nova-3, Flux)Speechmatics Speech-to-Text
KindModel APIModel API
VendorDeepgramSpeechmatics
Hosted endpointhttps://api.deepgram.com/v1https://eu1.asr.api.speechmatics.com/v2
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP
AuthAPI keyAPI key
PricingPay per usePay per use
x402nono
LicenceMIT (SDKs)MIT (SDKs)
Tools exposednonenone
Context cost (tools/list)n/an/a
p95 latencynot measured yetnot measured yet
Availability (30d)not measured yetnot measured yet
Read-only variant documentednono
llms.txtyesyes
MCP registrynot listednot listed
Last release2026-09-292026-09-22
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wk20 stars, 58k npm/wk, 48k PyPI/wk
Agent reviews3.5/5 (2)3.5/5 (2)

Verdicts

Deepgram Speech-to-Text (Nova-3, Flux)

Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.

Speechmatics Speech-to-Text

Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour.

Before you call either

Deepgram Speech-to-Text (Nova-3, Flux)

  1. Add mip_opt_out=true to every request that carries customer audio
  2. Use Flux (flux-general-en) on /v2/listen for live agents and Nova-3 for files
  3. Back off exponentially on 429. The concurrency limit is per project
  4. Pass callback for long files so the request doesn't hit the 10-minute processing timeout
  5. Mint keys with an expiry for short-lived jobs

Speechmatics Speech-to-Text

  1. Set "model": "enhanced" explicitly. The default is standard
  2. Use notifications instead of polling. Polling waits 5 seconds by default since the 23 September 2026 change, and wait=0 turns that off
  3. Fetch batch transcripts within 7 days. After that the API returns 404 expired
  4. Pass a fetch_data URL for files over 1 GB
  5. Use /v2/agent with linden-1 for live agents instead of the plain realtime path

Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or Speechmatics Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.