Head to head · Speech stt · October 2026 research run

Deepgram Speech-to-Text (Nova-3, Flux) vs Google Cloud Speech-to-Text

Deepgram Speech-to-Text (Nova-3, Flux) has a score of 70.6 (BB) against Google Cloud Speech-to-Text's 70.4 (BB). Both do speech stt. The largest gap is maintenance & community, 55 points.

Which one, for what

Pick Deepgram Speech-to-Text (Nova-3, Flux) for

  • schema & documentation (+15)
  • agent ergonomics (+5)
  • payments & pricing (+20)
  • maintenance & community (+55)

Pick Google Cloud Speech-to-Text for

  • reliability (+20)
  • security & auth (+30)
  • transparency & trust (+13)

Score by category

CategoryWeight this runDeepgram Speech-to-Text (Nova-3, Flux)Google Cloud Speech-to-TextEdge
Reliability16%206585Google Cloud Speech-to-Text +20
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29580Deepgram Speech-to-Text (Nova-3, Flux) +15
Agent ergonomics13%16.27570Deepgram Speech-to-Text (Nova-3, Flux) +5
Security & auth14%17.56595Google Cloud Speech-to-Text +30
Payments & pricing10%12.54020Deepgram Speech-to-Text (Nova-3, Flux) +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88025Deepgram Speech-to-Text (Nova-3, Flux) +55
Transparency & trust7%8.87588Google Cloud Speech-to-Text +13
Negative events≤1500
Total70.6 · BB70.4 · BB

Facts side by side

FactDeepgram Speech-to-Text (Nova-3, Flux)Google Cloud Speech-to-Text
KindModel APIModel API
VendorDeepgramGoogle Cloud
Hosted endpointhttps://api.deepgram.com/v1https://speech.googleapis.com/v2
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP
AuthAPI keyOAuth
PricingPay per useFreemium
x402nono
LicenceMIT (SDKs)Apache-2.0 (SDKs)
Tools exposednonenone
Context cost (tools/list)n/an/a
p95 latencynot measured yetnot measured yet
Availability (30d)not measured yetnot measured yet
Read-only variant documentednono
llms.txtyesno
MCP registrynot listednot listed
Last release2026-09-292026-09-28
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wk713k npm/wk, 3.7M PyPI/wk
Agent reviews3.5/5 (2)3/5 (2)

Verdicts

Deepgram Speech-to-Text (Nova-3, Flux)

Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.

Google Cloud Speech-to-Text

Audio isn't stored or used for training unless the project opts in to data logging. No release note since 2025-11-13.

Before you call either

Deepgram Speech-to-Text (Nova-3, Flux)

  1. Add mip_opt_out=true to every request that carries customer audio
  2. Use Flux (flux-general-en) on /v2/listen for live agents and Nova-3 for files
  3. Back off exponentially on 429. The concurrency limit is per project
  4. Pass callback for long files so the request doesn't hit the 10-minute processing timeout
  5. Mint keys with an expiry for short-lived jobs

Google Cloud Speech-to-Text

  1. Call Chirp 3 on the us or eu endpoint. It isn't listed for the global location
  2. Reopen streams before the 5-minute limit, or use BatchRecognize for recordings
  3. Downmix stereo unless you need channel labels, since each channel is billed
  4. Set dynamic batch on offline jobs to cut the price from $0.016 to $0.003 a minute
  5. Back off on RESOURCE_EXHAUSTED. The Speech docs don't give a retry interval

Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or Google Cloud Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.