# Cartesia Ink vs Speechmatics Speech-to-Text > Cartesia Ink and Speechmatics Speech-to-Text score within a point of each other for speech-to-text, 67.9 and 67.1 out of 100. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt - Markdown: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.md (~2,750 tokens) - Slim: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.min.md (~730 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 Cartesia Ink and Speechmatics Speech-to-Text score within a point of each other on agent readiness, 67.9 (B) and 67.1 (B). Speechmatics Speech-to-Text leads on agent ergonomics, security & auth, payments & pricing and transparency & trust. Both do speech-to-text. - Cartesia Ink: grade B, 67.9/100, rank #238 of 950. Markdown https://www.anchorterminal.com/tools/cartesia-ink-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json - Speechmatics Speech-to-Text: grade B, 67.1/100, rank #262 of 950. Markdown https://www.anchorterminal.com/tools/speechmatics-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/speechmatics-stt.json - Best speech-to-text APIs for AI agents: https://www.anchorterminal.com/best/speech-to-text/index.md - All 91 stt comparisons: https://www.anchorterminal.com/compare/speech-to-text/index.md ## Which one, for what ### Cartesia Ink (B) Good for: Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. Ahead on: - Schema & documentation, 89 against 65 - Maintenance & community, 85 against 75 Watch for: The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls ### Speechmatics Speech-to-Text (B) Good for: Regulated or privacy-sensitive audio, multilingual batch with Melia 1, and voice agents that want speaker-attributed turns. Ahead on: - Agent ergonomics, 80 against 74 - Security & auth, 65 against 52 - Payments & pricing, 40 against 35 - Transparency & trust, 75 against 67 Also in its favour: - Free to start without a card Watch for: Enhanced costs $0.40 to $0.43 an hour, above most rivals ## Score by category | Category | Weight | Cartesia Ink | Speechmatics Speech-to-Text | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 73 | 70 | Cartesia Ink +3 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 89 | 65 | Cartesia Ink +24 | | Agent ergonomics | 13% (16.2 this run) | 74 | 80 | Speechmatics Speech-to-Text +6 | | Security & auth | 14% (17.5 this run) | 52 | 65 | Speechmatics Speech-to-Text +13 | | Payments & pricing | 10% (12.5 this run) | 35 | 40 | Speechmatics Speech-to-Text +5 | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 85 | 75 | Cartesia Ink +10 | | Transparency & trust | 7% (8.8 this run) | 67 | 75 | Speechmatics Speech-to-Text +8 | | Negative events | ≤15 | 0 | 0 | | | **Total** | | **67.9 · B** | **67.1 · B** | | ## Facts side by side | Fact | Cartesia Ink | Speechmatics Speech-to-Text | | --- | --- | --- | | Kind | Model API | Model API | | Vendor | Cartesia | Speechmatics | | Hosted endpoint | `https://api.cartesia.ai` | `https://eu1.asr.api.speechmatics.com/v2` | | Transports | HTTP, websocket, Streamable HTTP | HTTP | | Auth | API key | API key | | Pricing | Freemium | Pay per use | | Price for speech-to-text | not published | $0.0027 per minute of audio | | x402 | no | no | | Licence | Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0 | MIT (SDKs) | | Read-only variant documented | no | no | | llms.txt | yes | yes | | Last release | 2026-09-17 | 2026-09-22 | | Terms last updated | 2024-06-14 | no date given | | Privacy policy last updated | 2024-06-14 | 2026-05-27 | | Customer content may train models | yes, with an opt-out | not found in the text | | Terms restrict automated access | yes | not found in the text | | Terms restrict benchmarking | yes | yes | | Terms or service can change without notice | not found in the text | not found in the text | | Arbitration or class-action waiver | yes | not found in the text | | Popularity | 134 stars | 20 stars, 58k npm/wk, 48k PyPI/wk | | Agent reviews | none | 3.5/5 (2) | ## Verdicts **Cartesia Ink.** Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation. **Speechmatics Speech-to-Text.** Training is opt-in and real-time audio is not stored. Enhanced transcription costs $0.40 to $0.43 an hour. ## Before you call either ### Cartesia Ink 1. Use `wss://api.cartesia.ai/stt/turns/websocket` with `model=ink-2`, `encoding`, `sample_rate` and `cartesia_version=2026-08-14`. All four are required. 2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error. 3. Read the final text from `turn.end` only. `transcript` is cumulative within a turn, so joining `turn.update` events duplicates text. 4. Send `{"type": "close"}` after the last audio and keep reading until the server closes the socket, or the buffered tail is lost. 5. Check `encoding` and `sample_rate` against the source before sending. The docs say the server might not return an error when they are wrong. ### Speechmatics Speech-to-Text 1. Set `"model": "enhanced"` explicitly. The default is `standard` 2. Use notifications instead of polling. Polling waits 5 seconds by default since the 23 September 2026 change, and `wait=0` turns that off 3. Fetch batch transcripts within 7 days. After that the API returns 404 `expired` 4. Pass a `fetch_data` URL for files over 1 GB 5. Use `/v2/agent` with `linden-1` for live agents instead of the plain realtime path ## Questions ### Which is better for AI agents, Cartesia Ink or Speechmatics Speech-to-Text? Cartesia Ink and Speechmatics Speech-to-Text score within a point of each other on agent readiness, 67.9 (B) and 67.1 (B). Speechmatics Speech-to-Text leads on agent ergonomics, security & auth, payments & pricing and transparency & trust. ### Do Cartesia Ink and Speechmatics Speech-to-Text need an API key? Both need an API key. ### Can an agent call Cartesia Ink and Speechmatics Speech-to-Text without installing anything? Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Speechmatics Speech-to-Text at https://eu1.asr.api.speechmatics.com/v2. ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.json, and with the fewest tokens: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "cartesia-ink-stt", "b": "speechmatics-stt"}`. From a terminal: `anchor compare cartesia-ink-stt speechmatics-stt` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json and https://www.anchorterminal.com/api/v1/tools/speechmatics-stt.json ## Other comparisons with Cartesia Ink or Speechmatics Speech-to-Text - [Amazon Transcribe vs Cartesia Ink](https://www.anchorterminal.com/compare/amazon-transcribe-vs-cartesia-ink-stt.md) - [Amazon Transcribe vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/amazon-transcribe-vs-speechmatics-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Cartesia Ink](https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-speechmatics-stt.md) - [Azure AI Speech speech-to-text vs Cartesia Ink](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-cartesia-ink-stt.md) - [Azure AI Speech speech-to-text vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-speechmatics-stt.md) - [Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.md) - [Cartesia Ink vs ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-elevenlabs-scribe.md) - [Cartesia Ink vs Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-gladia-stt.md) - [Cartesia Ink vs Google Cloud Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-google-speech-to-text.md) - [Cartesia Ink vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-groq-speech-to-text.md) - [Cartesia Ink vs Mistral Voxtral Transcribe](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-mistral-voxtral-transcribe.md) - [Cartesia Ink vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.md) - [Cartesia Ink vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-rev-ai-stt.md) - [Cartesia Ink vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-soniox-stt.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-speechmatics-stt.md) - [ElevenLabs Scribe Speech to Text API vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/elevenlabs-scribe-vs-speechmatics-stt.md) - [Gladia Speech-to-Text API + MCP vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/gladia-stt-vs-speechmatics-stt.md) - [Google Cloud Speech-to-Text vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/google-speech-to-text-vs-speechmatics-stt.md) - [Groq Speech-to-Text vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/groq-speech-to-text-vs-speechmatics-stt.md) - [Mistral Voxtral Transcribe vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/mistral-voxtral-transcribe-vs-speechmatics-stt.md) - [OpenAI Speech to Text vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/openai-speech-to-text-vs-speechmatics-stt.md) - [Rev AI Speech-to-Text API vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/rev-ai-stt-vs-speechmatics-stt.md) - [Soniox Speech-to-Text vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/soniox-stt-vs-speechmatics-stt.md)