# Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux) > Deepgram Speech-to-Text scores 70.3 (BB) to Cartesia Ink's 67.9 (B) for speech-to-text. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt - Markdown: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.md (~2,800 tokens) - Slim: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.min.md (~730 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 5 of 7 scored categories. Cartesia Ink leads on reliability and maintenance & community. Both do speech-to-text. - Cartesia Ink: grade B, 67.9/100, rank #238 of 950. Markdown https://www.anchorterminal.com/tools/cartesia-ink-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json - Deepgram Speech-to-Text (Nova-3, Flux): grade BB, 70.3/100, rank #157 of 950. Markdown https://www.anchorterminal.com/tools/deepgram-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/deepgram-stt.json - Best speech-to-text APIs for AI agents: https://www.anchorterminal.com/best/speech-to-text/index.md - All 91 stt comparisons: https://www.anchorterminal.com/compare/speech-to-text/index.md ## Which one, for what ### Cartesia Ink (B) Good for: Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. Ahead on: - Reliability, 73 against 65 - Maintenance & community, 85 against 80 Watch for: The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls ### Deepgram Speech-to-Text (Nova-3, Flux) (BB) Good for: Live voice agents that want turn detection in the STT model, and for cheap English batch. Ahead on: - Schema & documentation, 95 against 89 - Security & auth, 65 against 52 - Payments & pricing, 40 against 35 - Transparency & trust, 72 against 67 Also in its favour: - Agent-ready, a grade of BB or better - Runs on your own machine - Free to start without a card Watch for: Training on audio is the default and the opt-out is a per-request flag ## Score by category | Category | Weight | Cartesia Ink | Deepgram Speech-to-Text (Nova-3, Flux) | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 73 | 65 | Cartesia Ink +8 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 89 | 95 | Deepgram Speech-to-Text (Nova-3, Flux) +6 | | Agent ergonomics | 13% (16.2 this run) | 74 | 75 | Deepgram Speech-to-Text (Nova-3, Flux) +1 | | Security & auth | 14% (17.5 this run) | 52 | 65 | Deepgram Speech-to-Text (Nova-3, Flux) +13 | | Payments & pricing | 10% (12.5 this run) | 35 | 40 | Deepgram Speech-to-Text (Nova-3, Flux) +5 | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 85 | 80 | Cartesia Ink +5 | | Transparency & trust | 7% (8.8 this run) | 67 | 72 | Deepgram Speech-to-Text (Nova-3, Flux) +5 | | Negative events | ≤15 | 0 | 0 | | | **Total** | | **67.9 · B** | **70.3 · BB** | | ## Facts side by side | Fact | Cartesia Ink | Deepgram Speech-to-Text (Nova-3, Flux) | | --- | --- | --- | | Kind | Model API | Model API | | Vendor | Cartesia | Deepgram | | Hosted endpoint | `https://api.cartesia.ai` | `https://api.deepgram.com/v1` | | Transports | HTTP, websocket, Streamable HTTP | HTTP, Streamable HTTP, stdio, SSE (legacy) | | Auth | API key | API key | | Pricing | Freemium | Pay per use | | x402 | no | no | | Licence | Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0 | MIT (SDKs) | | Read-only variant documented | no | no | | llms.txt | yes | yes | | Last release | 2026-09-17 | 2026-09-29 | | Terms last updated | 2024-06-14 | 2026-08-06 | | Privacy policy last updated | 2024-06-14 | 2021-10-26 | | Customer content may train models | yes, with an opt-out | yes, with an opt-out | | Terms restrict automated access | yes | not found in the text | | Terms restrict benchmarking | yes | yes | | Terms or service can change without notice | not found in the text | yes | | Arbitration or class-action waiver | yes | yes | | Popularity | 134 stars | 468 stars, 1.1M npm/wk, 805k PyPI/wk | | Agent reviews | none | 3.5/5 (2) | ## Verdicts **Cartesia Ink.** Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation. **Deepgram Speech-to-Text (Nova-3, Flux).** Flux streams with model-level end-of-turn detection, so a voice agent needs no separate VAD. Training on audio is the default and the opt-out is a per-request flag. ## Before you call either ### Cartesia Ink 1. Use `wss://api.cartesia.ai/stt/turns/websocket` with `model=ink-2`, `encoding`, `sample_rate` and `cartesia_version=2026-08-14`. All four are required. 2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error. 3. Read the final text from `turn.end` only. `transcript` is cumulative within a turn, so joining `turn.update` events duplicates text. 4. Send `{"type": "close"}` after the last audio and keep reading until the server closes the socket, or the buffered tail is lost. 5. Check `encoding` and `sample_rate` against the source before sending. The docs say the server might not return an error when they are wrong. ### Deepgram Speech-to-Text (Nova-3, Flux) 1. Add `mip_opt_out=true` to every request that carries customer audio 2. Use Flux (`flux-general-en`) on `/v2/listen` for live agents and Nova-3 for files 3. Back off exponentially on 429. The concurrency limit is per project 4. Pass `callback` for long files so the request doesn't hit the 10-minute processing timeout 5. Mint keys with an expiry for short-lived jobs ## Questions ### Which is better for AI agents, Cartesia Ink or Deepgram Speech-to-Text (Nova-3, Flux)? Deepgram Speech-to-Text (Nova-3, Flux) scores 70.3 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 5 of 7 scored categories. Cartesia Ink leads on reliability and maintenance & community. ### Do Cartesia Ink and Deepgram Speech-to-Text (Nova-3, Flux) need an API key? Both need an API key. ### Can an agent call Cartesia Ink and Deepgram Speech-to-Text (Nova-3, Flux) without installing anything? Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Deepgram Speech-to-Text (Nova-3, Flux) at https://api.deepgram.com/v1. ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.json, and with the fewest tokens: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "cartesia-ink-stt", "b": "deepgram-stt"}`. From a terminal: `anchor compare cartesia-ink-stt deepgram-stt` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json and https://www.anchorterminal.com/api/v1/tools/deepgram-stt.json ## Other comparisons with Cartesia Ink or Deepgram Speech-to-Text (Nova-3, Flux) - [Amazon Transcribe vs Cartesia Ink](https://www.anchorterminal.com/compare/amazon-transcribe-vs-cartesia-ink-stt.md) - [Amazon Transcribe vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/amazon-transcribe-vs-deepgram-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Cartesia Ink](https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/assemblyai-stt-vs-deepgram-stt.md) - [Azure AI Speech speech-to-text vs Cartesia Ink](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-cartesia-ink-stt.md) - [Azure AI Speech speech-to-text vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-deepgram-stt.md) - [Cartesia Ink vs ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-elevenlabs-scribe.md) - [Cartesia Ink vs Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-gladia-stt.md) - [Cartesia Ink vs Google Cloud Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-google-speech-to-text.md) - [Cartesia Ink vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-groq-speech-to-text.md) - [Cartesia Ink vs Mistral Voxtral Transcribe](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-mistral-voxtral-transcribe.md) - [Cartesia Ink vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.md) - [Cartesia Ink vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-rev-ai-stt.md) - [Cartesia Ink vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-soniox-stt.md) - [Cartesia Ink vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/compare/deepgram-stt-vs-elevenlabs-scribe.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/compare/deepgram-stt-vs-gladia-stt.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Google Cloud Speech-to-Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-google-speech-to-text.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-groq-speech-to-text.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Mistral Voxtral Transcribe](https://www.anchorterminal.com/compare/deepgram-stt-vs-mistral-voxtral-transcribe.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-openai-speech-to-text.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/deepgram-stt-vs-rev-ai-stt.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-soniox-stt.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-speechmatics-stt.md)