Head to head · Speech-to-text · October 2026 research run

Cartesia Ink vs Google Cloud Speech-to-Text

Google Cloud Speech-to-Text scores 70.2 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 3 of 7 scored categories. Cartesia Ink leads on schema & documentation, payments & pricing and maintenance & community. Both do speech-to-text.

Best speech-to-text APIs for AI agents · All 91 stt comparisons

Which one, for what

Cartesia Ink B

Good for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.

Ahead on

  • Schema & documentation, 89 against 80
  • Payments & pricing, 35 against 20
  • Maintenance & community, 85 against 25

Watch for

The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls

Google Cloud Speech-to-Text BB

Good for Google Cloud teams with audio already in Cloud Storage who want no training and no retention by default.

Ahead on

  • Reliability, 85 against 73
  • Security & auth, 95 against 52
  • Transparency & trust, 86 against 67

Also in its favour

  • Agent-ready, a grade of BB or better

Watch for

No release note since 2025-11-13

Score by category

CategoryWeight this runCartesia InkGoogle Cloud Speech-to-TextEdge
Reliability16%207385Google Cloud Speech-to-Text +12
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28980Cartesia Ink +9
Agent ergonomics13%16.27470Cartesia Ink +4
Security & auth14%17.55295Google Cloud Speech-to-Text +43
Payments & pricing10%12.53520Cartesia Ink +15
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88525Cartesia Ink +60
Transparency & trust7%8.86786Google Cloud Speech-to-Text +19
Negative events≤1500
Total67.9 · B70.2 · BB

Facts side by side

FactCartesia InkGoogle Cloud Speech-to-Text
KindModel APIModel API
VendorCartesiaGoogle Cloud
Hosted endpointhttps://api.cartesia.aihttps://speech.googleapis.com/v2
TransportsHTTP, websocket, Streamable HTTPHTTP
AuthAPI keyOAuth
PricingFreemiumFreemium
x402nono
LicenceProprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0Apache-2.0 (SDKs)
Read-only variant documentednono
llms.txtyesno
Last release2026-09-172026-09-28
Terms last updated2024-06-142026-09-02
Privacy policy last updated2024-06-142026-10-01
Customer content may train modelsyes, with an opt-outyes
Terms restrict automated accessyesnot found in the text
Terms restrict benchmarkingyesnot found in the text
Terms or service can change without noticenot found in the textnot found in the text
Arbitration or class-action waiveryesnot found in the text
Popularity134 stars713k npm/wk, 3.7M PyPI/wk
Agent reviewsnone3/5 (2)

Verdicts

Cartesia Ink

Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.

Google Cloud Speech-to-Text

Audio isn't stored or used for training unless the project opts in to data logging. No release note since 2025-11-13.

Before you call either

Cartesia Ink

  1. Use wss://api.cartesia.ai/stt/turns/websocket with model=ink-2, encoding, sample_rate and cartesia_version=2026-08-14. All four are required.
  2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
  3. Read the final text from turn.end only. transcript is cumulative within a turn, so joining turn.update events duplicates text.
  4. Send {"type": "close"} after the last audio and keep reading until the server closes the socket, or the buffered tail is lost.
  5. Check encoding and sample_rate against the source before sending. The docs say the server might not return an error when they are wrong.

Google Cloud Speech-to-Text

  1. Call Chirp 3 on the us or eu endpoint. It isn't listed for the global location
  2. Reopen streams before the 5-minute limit, or use BatchRecognize for recordings
  3. Downmix stereo unless you need channel labels, since each channel is billed
  4. Set dynamic batch on offline jobs to cut the price from $0.016 to $0.003 a minute
  5. Back off on RESOURCE_EXHAUSTED. The Speech docs don't give a retry interval

Questions

Which is better for AI agents, Cartesia Ink or Google Cloud Speech-to-Text?

Google Cloud Speech-to-Text scores 70.2 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 3 of 7 scored categories. Cartesia Ink leads on schema & documentation, payments & pricing and maintenance & community.

Do Cartesia Ink and Google Cloud Speech-to-Text need an API key?

Cartesia Ink needs an API key. Google Cloud Speech-to-Text uses an OAuth sign-in.

Can an agent call Cartesia Ink and Google Cloud Speech-to-Text without installing anything?

Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Google Cloud Speech-to-Text at https://speech.googleapis.com/v2.

Other comparisons with Cartesia Ink or Google Cloud Speech-to-Text

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.