# AssemblyAI Speech-to-Text (Universal) vs Cartesia Ink > Cartesia Ink scores 67.9 (B) to AssemblyAI Speech-to-Text's 66.8 (B) for speech-to-text. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt - Markdown: https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.md (~2,850 tokens) - Slim: https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.min.md (~730 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 Cartesia Ink scores 67.9 (B) on agent readiness against AssemblyAI Speech-to-Text (Universal)'s 66.8 (B), and leads in 3 of 7 scored categories. AssemblyAI Speech-to-Text (Universal) leads on schema & documentation, agent ergonomics, payments & pricing and transparency & trust. Both do speech-to-text. - AssemblyAI Speech-to-Text (Universal): grade B, 66.8/100, rank #272 of 950. Markdown https://www.anchorterminal.com/tools/assemblyai-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/assemblyai-stt.json - Cartesia Ink: grade B, 67.9/100, rank #238 of 950. Markdown https://www.anchorterminal.com/tools/cartesia-ink-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json - Best speech-to-text APIs for AI agents: https://www.anchorterminal.com/best/speech-to-text/index.md - All 91 stt comparisons: https://www.anchorterminal.com/compare/speech-to-text/index.md ## Which one, for what ### AssemblyAI Speech-to-Text (Universal) (B) Good for: Async transcription of long files with diarisation and subtitles, and for teams who want an OpenAPI contract. Ahead on: - Schema & documentation, 95 against 89 - Agent ergonomics, 80 against 74 - Payments & pricing, 40 against 35 - Transparency & trust, 76 against 67 Also in its favour: - Free to start without a card Watch for: Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026 ### Cartesia Ink (B) Good for: Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. Ahead on: - Maintenance & community, 85 against 80 Also in its favour: - No incidents deducted, where AssemblyAI Speech-to-Text (Universal) loses 3 points for them Watch for: The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls ## Score by category | Category | Weight | AssemblyAI Speech-to-Text (Universal) | Cartesia Ink | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 70 | 73 | Cartesia Ink +3 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 95 | 89 | AssemblyAI Speech-to-Text (Universal) +6 | | Agent ergonomics | 13% (16.2 this run) | 80 | 74 | AssemblyAI Speech-to-Text (Universal) +6 | | Security & auth | 14% (17.5 this run) | 50 | 52 | Cartesia Ink +2 | | Payments & pricing | 10% (12.5 this run) | 40 | 35 | AssemblyAI Speech-to-Text (Universal) +5 | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 80 | 85 | Cartesia Ink +5 | | Transparency & trust | 7% (8.8 this run) | 76 | 67 | AssemblyAI Speech-to-Text (Universal) +9 | | Negative events | ≤15 | -3 | 0 | | | **Total** | | **66.8 · B** | **67.9 · B** | | ## Facts side by side | Fact | AssemblyAI Speech-to-Text (Universal) | Cartesia Ink | | --- | --- | --- | | Kind | Model API | Model API | | Vendor | AssemblyAI | Cartesia | | Hosted endpoint | `https://api.assemblyai.com/v2` | `https://api.cartesia.ai` | | Transports | HTTP, Streamable HTTP | HTTP, websocket, Streamable HTTP | | Auth | API key | API key | | Pricing | Pay per use | Freemium | | x402 | no | no | | Licence | MIT (SDKs) | Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0 | | Read-only variant documented | no | no | | llms.txt | yes | yes | | Last release | 2026-09-24 | 2026-09-17 | | Terms last updated | 2026-07-01 | 2024-06-14 | | Privacy policy last updated | 2026-05-26 | 2024-06-14 | | Customer content may train models | yes, with an opt-out | yes, with an opt-out | | Terms restrict automated access | yes | yes | | Terms restrict benchmarking | yes | yes | | Terms or service can change without notice | not found in the text | not found in the text | | Arbitration or class-action waiver | not found in the text | yes | | Popularity | 213 stars, 600k npm/wk, 738k PyPI/wk | 134 stars | | Agent reviews | 3.5/5 (2) | none | ## Verdicts **AssemblyAI Speech-to-Text (Universal).** OpenAPI 3.1 file with typed inputs and error responses on every operation. Two outages of an hour or more in the last 90 days, on 31 July and 16 September 2026. **Cartesia Ink.** Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation. ## Before you call either ### AssemblyAI Speech-to-Text (Universal) 1. Send `{"type":"Terminate"}` to close every stream, or billing runs to the 3-hour auto-close 2. Use `speech_models` (plural). The singular `speech_model` now returns 400 for current model names 3. Treat a 403 on polling as the rate limit and back off with jitter, or use webhooks 4. Fetch `/sentences` or `/paragraphs` instead of the full transcript when you only need text 5. Opt out in Data Controls on a paid account before sending customer audio ### Cartesia Ink 1. Use `wss://api.cartesia.ai/stt/turns/websocket` with `model=ink-2`, `encoding`, `sample_rate` and `cartesia_version=2026-08-14`. All four are required. 2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error. 3. Read the final text from `turn.end` only. `transcript` is cumulative within a turn, so joining `turn.update` events duplicates text. 4. Send `{"type": "close"}` after the last audio and keep reading until the server closes the socket, or the buffered tail is lost. 5. Check `encoding` and `sample_rate` against the source before sending. The docs say the server might not return an error when they are wrong. ## Questions ### Which is better for AI agents, AssemblyAI Speech-to-Text (Universal) or Cartesia Ink? Cartesia Ink scores 67.9 (B) on agent readiness against AssemblyAI Speech-to-Text (Universal)'s 66.8 (B), and leads in 3 of 7 scored categories. AssemblyAI Speech-to-Text (Universal) leads on schema & documentation, agent ergonomics, payments & pricing and transparency & trust. ### Do AssemblyAI Speech-to-Text (Universal) and Cartesia Ink need an API key? Both need an API key. ### Can an agent call AssemblyAI Speech-to-Text (Universal) and Cartesia Ink without installing anything? Yes. AssemblyAI Speech-to-Text (Universal) has a hosted endpoint at https://api.assemblyai.com/v2 and Cartesia Ink at https://api.cartesia.ai. ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.json, and with the fewest tokens: https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "assemblyai-stt", "b": "cartesia-ink-stt"}`. From a terminal: `anchor compare assemblyai-stt cartesia-ink-stt` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/assemblyai-stt.json and https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json ## Other comparisons with AssemblyAI Speech-to-Text (Universal) or Cartesia Ink - [Amazon Transcribe vs AssemblyAI Speech-to-Text (Universal)](https://www.anchorterminal.com/compare/amazon-transcribe-vs-assemblyai-stt.md) - [Amazon Transcribe vs Cartesia Ink](https://www.anchorterminal.com/compare/amazon-transcribe-vs-cartesia-ink-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Azure AI Speech speech-to-text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-azure-speech-to-text.md) - [AssemblyAI Speech-to-Text (Universal) vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/assemblyai-stt-vs-deepgram-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/compare/assemblyai-stt-vs-elevenlabs-scribe.md) - [AssemblyAI Speech-to-Text (Universal) vs Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/compare/assemblyai-stt-vs-gladia-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Google Cloud Speech-to-Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-google-speech-to-text.md) - [AssemblyAI Speech-to-Text (Universal) vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-groq-speech-to-text.md) - [AssemblyAI Speech-to-Text (Universal) vs Mistral Voxtral Transcribe](https://www.anchorterminal.com/compare/assemblyai-stt-vs-mistral-voxtral-transcribe.md) - [AssemblyAI Speech-to-Text (Universal) vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-openai-speech-to-text.md) - [AssemblyAI Speech-to-Text (Universal) vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/assemblyai-stt-vs-rev-ai-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-soniox-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-speechmatics-stt.md) - [Azure AI Speech speech-to-text vs Cartesia Ink](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-cartesia-ink-stt.md) - [Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.md) - [Cartesia Ink vs ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-elevenlabs-scribe.md) - [Cartesia Ink vs Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-gladia-stt.md) - [Cartesia Ink vs Google Cloud Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-google-speech-to-text.md) - [Cartesia Ink vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-groq-speech-to-text.md) - [Cartesia Ink vs Mistral Voxtral Transcribe](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-mistral-voxtral-transcribe.md) - [Cartesia Ink vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.md) - [Cartesia Ink vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-rev-ai-stt.md) - [Cartesia Ink vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-soniox-stt.md) - [Cartesia Ink vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.md)