Head to head · Speech-to-text · October 2026 research run

Cartesia Ink vs Speechmatics Speech-to-Text

Cartesia Ink and Speechmatics Speech-to-Text score within a point of each other on agent readiness, 67.9 (B) and 67.1 (B). Speechmatics Speech-to-Text leads on agent ergonomics, security & auth, payments & pricing and transparency & trust. Both do speech-to-text.

Best speech-to-text APIs for AI agents · All 91 stt comparisons

Which one, for what

Cartesia Ink B

Good for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.

Ahead on

  • Schema & documentation, 89 against 65
  • Maintenance & community, 85 against 75

Watch for

The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls

Speechmatics Speech-to-Text B

Good for Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns.

Ahead on

  • Agent ergonomics, 80 against 74
  • Security & auth, 65 against 52
  • Payments & pricing, 40 against 35
  • Transparency & trust, 75 against 67

Also in its favour

  • Free to start without a card

Watch for

Enhanced costs $0.40 to $0.43 an hour, above most rivals

Score by category

CategoryWeight this runCartesia InkSpeechmatics Speech-to-TextEdge
Reliability16%207370Cartesia Ink +3
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28965Cartesia Ink +24
Agent ergonomics13%16.27480Speechmatics Speech-to-Text +6
Security & auth14%17.55265Speechmatics Speech-to-Text +13
Payments & pricing10%12.53540Speechmatics Speech-to-Text +5
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88575Cartesia Ink +10
Transparency & trust7%8.86775Speechmatics Speech-to-Text +8
Negative events≤1500
Total67.9 · B67.1 · B

Facts side by side

FactCartesia InkSpeechmatics Speech-to-Text
KindModel APIModel API
VendorCartesiaSpeechmatics
Hosted endpointhttps://api.cartesia.aihttps://eu1.asr.api.speechmatics.com/v2
TransportsHTTP, websocket, Streamable HTTPHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
Price for speech-to-textnot published$0.0027 per minute of audio
x402nono
LicenceProprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0MIT (SDKs)
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-172026-09-22
Terms last updated2024-06-14no date given
Privacy policy last updated2024-06-142026-05-27
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waiveryesnot found in the text
Popularity134 stars20 stars, 58k npm/wk, 48k PyPI/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Cartesia Ink

Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.

Speechmatics Speech-to-Text

Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour.

Before you call either

Cartesia Ink

  1. Use wss://api.cartesia.ai/stt/turns/websocket with model=ink-2, encoding, sample_rate and cartesia_version=2026-08-14. All four are required.
  2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
  3. Read the final text from turn.end only. transcript is cumulative within a turn, so joining turn.update events duplicates text.
  4. Send {"type": "close"} after the last audio and keep reading until the server closes the socket, or the buffered tail is lost.
  5. Check encoding and sample_rate against the source before sending. The docs say the server might not return an error when they are wrong.

Speechmatics Speech-to-Text

  1. Set "model": "enhanced" explicitly. The default is standard
  2. Use notifications instead of polling. Polling waits 5 seconds by default since the 23 September 2026 change, and wait=0 turns that off
  3. Fetch batch transcripts within 7 days. After that the API returns 404 expired
  4. Pass a fetch_data URL for files over 1 GB
  5. Use /v2/agent with linden-1 for live agents instead of the plain realtime path

Questions

Which is better for AI agents, Cartesia Ink or Speechmatics Speech-to-Text?

Cartesia Ink and Speechmatics Speech-to-Text score within a point of each other on agent readiness, 67.9 (B) and 67.1 (B). Speechmatics Speech-to-Text leads on agent ergonomics, security & auth, payments & pricing and transparency & trust.

Do Cartesia Ink and Speechmatics Speech-to-Text need an API key?

Both need an API key.

Can an agent call Cartesia Ink and Speechmatics Speech-to-Text without installing anything?

Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Speechmatics Speech-to-Text at https://eu1.asr.api.speechmatics.com/v2.

Other comparisons with Cartesia Ink or Speechmatics Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.