Head to head · Speech stt · October 2026 research run

Deepgram Speech-to-Text (Nova-3, Flux) vs Gladia Speech-to-Text API + MCP

Deepgram Speech-to-Text (Nova-3, Flux) has a score of 70.6 (BB) against Gladia Speech-to-Text API + MCP's 69.7 (B). Both do speech stt. The largest gap is transparency & trust, 9 points.

Which one, for what

Pick Deepgram Speech-to-Text (Nova-3, Flux) for

  • reliability (+5)
  • transparency & trust (+9)

Pick Gladia Speech-to-Text API + MCP for

  • security & auth (+5)

Score by category

CategoryWeight this runDeepgram Speech-to-Text (Nova-3, Flux)Gladia Speech-to-Text API + MCPEdge
Reliability16%206560Deepgram Speech-to-Text (Nova-3, Flux) +5
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.29595even
Agent ergonomics13%16.27575even
Security & auth14%17.56570Gladia Speech-to-Text API + MCP +5
Payments & pricing10%12.54040even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88080even
Transparency & trust7%8.87566Deepgram Speech-to-Text (Nova-3, Flux) +9
Negative events≤1500
Total70.6 · BB69.7 · B

Facts side by side

FactDeepgram Speech-to-Text (Nova-3, Flux)Gladia Speech-to-Text API + MCP
KindModel APIModel API
VendorDeepgramGladia
Hosted endpointhttps://api.deepgram.com/v1https://api.gladia.io/v2
TransportsHTTP, Streamable HTTP, stdio, SSE (legacy)HTTP, stdio
AuthAPI keyAPI key
PricingPay per useFreemium
x402nono
LicenceMIT (SDKs)MIT (SDKs and MCP server)
Tools exposednone8
Context cost (tools/list)n/an/a
p95 latencynot measured yetnot measured yet
Availability (30d)not measured yetnot measured yet
Read-only variant documentednono
llms.txtyesyes
MCP registrynot listednot listed
Last release2026-09-292026-09-24
Popularity468 stars, 1.1M npm/wk, 805k PyPI/wk4 stars, 4.8k npm/wk, 126k PyPI/wk
Agent reviews3.5/5 (2)3/5 (2)

Verdicts

Deepgram Speech-to-Text (Nova-3, Flux)

Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.

Gladia Speech-to-Text API + MCP

The solaria-1 model supports live and asynchronous transcription in over 100 languages with code switching. Starter pricing is $0.61 an hour for asynchronous transcription and $0.75 for real-time audio.

Before you call either

Deepgram Speech-to-Text (Nova-3, Flux)

  1. Add mip_opt_out=true to every request that carries customer audio
  2. Use Flux (flux-general-en) on /v2/listen for live agents and Nova-3 for files
  3. Back off exponentially on 429. The concurrency limit is per project
  4. Pass callback for long files so the request doesn't hit the 10-minute processing timeout
  5. Mint keys with an expiry for short-lived jobs

Gladia Speech-to-Text API + MCP

  1. Don't resubmit a pre-recorded job after a 200 or a transcription.created webhook. It's already queued
  2. Pick solaria-3 only for async EN, FR, DE, ES or IT audio. Anything live or multilingual needs solaria-1
  3. A 429 means the concurrency limit, 3 async and 1 live on the free plan. Wait for a running job to finish
  4. Upgrade off the free plan before sending sensitive audio
  5. Split files over 135 minutes or 1,000 MB

Other comparisons with Deepgram Speech-to-Text (Nova-3, Flux) or Gladia Speech-to-Text API + MCP

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.