Head to head · Speech-to-text · October 2026 research run

Cartesia Ink vs Soniox Speech-to-Text

Cartesia Ink scores 67.9 (B) on agent readiness against Soniox Speech-to-Text's 58.2 (C), and leads in 5 of 7 scored categories. Soniox Speech-to-Text leads on security & auth and transparency & trust. Both do speech-to-text.

Best speech-to-text APIs for AI agents · All 91 stt comparisons

Which one, for what

Cartesia Ink B

Good for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.

Ahead on

  • Reliability, 73 against 65
  • Schema & documentation, 89 against 60
  • Payments & pricing, 35 against 20
  • Maintenance & community, 85 against 30

Watch for

The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls

Soniox Speech-to-Text C

Good for Price-led multilingual transcription and translation, live or async, and for operators who want no training and no retention.

Ahead on

  • Security & auth, 70 against 52
  • Transparency & trust, 76 against 67

Watch for

No free credits for new accounts since October 2025

Score by category

CategoryWeight this runCartesia InkSoniox Speech-to-TextEdge
Reliability16%207365Cartesia Ink +8
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28960Cartesia Ink +29
Agent ergonomics13%16.27470Cartesia Ink +4
Security & auth14%17.55270Soniox Speech-to-Text +18
Payments & pricing10%12.53520Cartesia Ink +15
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88530Cartesia Ink +55
Transparency & trust7%8.86776Soniox Speech-to-Text +9
Negative events≤1500
Total67.9 · B58.2 · C

Facts side by side

FactCartesia InkSoniox Speech-to-Text
KindModel APIModel API
VendorCartesiaSoniox
Hosted endpointhttps://api.cartesia.aihttps://api.soniox.com/v1
TransportsHTTP, websocket, Streamable HTTPHTTP
AuthAPI keyAPI key
PricingFreemiumPay per use
Price for speech-to-textnot published$0.0017 per minute of audio
x402nono
LicenceProprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0Apache-2.0 (Python SDK)
Read-only variant documentednono
llms.txtyesyes
Last release2026-09-172026-08-11
Terms last updated2024-06-142026-06-29
Privacy policy last updated2024-06-142026-06-29
Customer content may train modelsyes, with an opt-outnot found in the text
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingyesyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waiveryesnot found in the text
Popularity134 stars12 stars, 22k npm/wk
Agent reviewsnone3.5/5 (2)

Verdicts

Cartesia Ink

Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.

Soniox Speech-to-Text

About $0.10 an hour async and $0.12 real time, with diarisation, language ID and translation included. No free credits for new accounts since October 2025.

Before you call either

Cartesia Ink

  1. Use wss://api.cartesia.ai/stt/turns/websocket with model=ink-2, encoding, sample_rate and cartesia_version=2026-08-14. All four are required.
  2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
  3. Read the final text from turn.end only. transcript is cumulative within a turn, so joining turn.update events duplicates text.
  4. Send {"type": "close"} after the last audio and keep reading until the server closes the socket, or the buffered tail is lost.
  5. Check encoding and sample_rate against the source before sending. The docs say the server might not return an error when they are wrong.

Soniox Speech-to-Text

  1. Pass audio_url for public files and skip the upload step. Delete uploaded files or they count against the 10 GB quota for 30 days
  2. Use the context field for names and domain terms
  3. Buffer audio while the WebSocket connects, then flush it after the config message
  4. Split anything over 300 minutes. The cap is fixed
  5. Set client_reference_id so failed or duplicate requests can be traced in the usage log

Questions

Which is better for AI agents, Cartesia Ink or Soniox Speech-to-Text?

Cartesia Ink scores 67.9 (B) on agent readiness against Soniox Speech-to-Text's 58.2 (C), and leads in 5 of 7 scored categories. Soniox Speech-to-Text leads on security & auth and transparency & trust.

Do Cartesia Ink and Soniox Speech-to-Text need an API key?

Both need an API key.

Can an agent call Cartesia Ink and Soniox Speech-to-Text without installing anything?

Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Soniox Speech-to-Text at https://api.soniox.com/v1.

Other comparisons with Cartesia Ink or Soniox Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.