Head to head · Speech-to-text · October 2026 research run

Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux)

Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 5 of 7 scored categories. Cartesia Ink leads on reliability and maintenance & community. Both do speech-to-text.

Best speech-to-text APIs for AI agents · All 91 stt comparisons

Which one, for what

Cartesia Ink B

Good for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.

Ahead on

  • Reliability, 73 against 65
  • Maintenance & community, 85 against 80

Watch for

The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls

Deepgram Speech-to-Text (Nova-3, Flux) BB

Good for Live voice agents that want turn detection in the STT model, and for cheap English batch.

Ahead on

  • Schema & documentation, 95 against 89
  • Security & auth, 65 against 52
  • Payments & pricing, 40 against 35
  • Transparency & trust, 72 against 67

Also in its favour

  • Agent-ready, a grade of BB or better
  • Runs on your own machine
  • Free to start without a card

Watch for

Training on audio is the default and the opt-out is a per-request flag

Score by category

CategoryWeight this runCartesia InkDeepgram Speech-to-Text (Nova-3, Flux)Edge
Reliability16%207365Cartesia Ink +8
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28995Deepgram Speech-to-Text (Nova-3, Flux) +6
Agent ergonomics13%16.27475Deepgram Speech-to-Text (Nova-3, Flux) +1
Security & auth14%17.55265Deepgram Speech-to-Text (Nova-3, Flux) +13
Payments & pricing10%12.53540Deepgram Speech-to-Text (Nova-3, Flux) +5
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88580Cartesia Ink +5
Transparency & trust7%8.86772Deepgram Speech-to-Text (Nova-3, Flux) +5
Negative events≤1500
Total67.9 · B70.3 · BB

Facts side by side

FactCartesia InkDeepgram Speech-to-Text (Nova-3, Flux)
KindModel APIModel API
VendorCartesiaDeepgram
Hosted endpointhttps://api.cartesia.aihttps://api.deepgram.com/v1
TransportsHTTP, websocket, Streamable HTTPHTTP, Streamable HTTP, stdio, SSE (legacy)
AuthAPI keyAPI key
PricingFreemiumPay per use
x402nono
LicenceProprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0MIT (SDKs)
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-172026-09-29
Terms last updated2024-06-142026-08-06
Privacy policy last updated2024-06-142021-10-26
Customer content may train modelsyes, with an opt-outyes, with an opt-out
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waiveryesyes
Popularity134 stars468 stars, 1.1M npm/wk, 805k PyPI/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Cartesia Ink

Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.

Deepgram Speech-to-Text (Nova-3, Flux)

Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag.

Before you call either

Cartesia Ink

  1. Use wss://api.cartesia.ai/stt/turns/websocket with model=ink-2, encoding, sample_rate and cartesia_version=2026-08-14. All four are required.
  2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
  3. Read the final text from turn.end only. transcript is cumulative within a turn, so joining turn.update events duplicates text.
  4. Send {"type": "close"} after the last audio and keep reading until the server closes the socket, or the buffered tail is lost.
  5. Check encoding and sample_rate against the source before sending. The docs say the server might not return an error when they are wrong.

Deepgram Speech-to-Text (Nova-3, Flux)

  1. Add mip_opt_out=true to every request that carries customer audio
  2. Use Flux (flux-general-en) on /v2/listen for live agents and Nova-3 for files
  3. Back off exponentially on 429. The concurrency limit is per project
  4. Pass callback for long files so the request doesn't hit the 10-minute processing timeout
  5. Mint keys with an expiry for short-lived jobs

Questions

Which is better for AI agents, Cartesia Ink or Deepgram Speech-to-Text (Nova-3, Flux)?

Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 5 of 7 scored categories. Cartesia Ink leads on reliability and maintenance & community.

Do Cartesia Ink and Deepgram Speech-to-Text (Nova-3, Flux) need an API key?

Both need an API key.

Can an agent call Cartesia Ink and Deepgram Speech-to-Text (Nova-3, Flux) without installing anything?

Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Deepgram Speech-to-Text (Nova-3, Flux) at https://api.deepgram.com/v1.

Other comparisons with Cartesia Ink or Deepgram Speech-to-Text (Nova-3, Flux)

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.