# Cartesia Ink vs OpenAI Speech to Text > OpenAI Speech to Text scores 72.4 (BB) to Cartesia Ink's 67.9 (B) for speech-to-text. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text - Markdown: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.md (~2,900 tokens) - Slim: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.min.md (~780 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 OpenAI Speech to Text scores 72.4 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 3 of 7 scored categories. Cartesia Ink leads on payments & pricing, maintenance & community and transparency & trust. Both do speech-to-text. - Cartesia Ink: grade B, 67.9/100, rank #238 of 950. Markdown https://www.anchorterminal.com/tools/cartesia-ink-stt.md · JSON https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json - OpenAI Speech to Text: grade BB, 72.4/100, rank #106 of 950. Markdown https://www.anchorterminal.com/tools/openai-speech-to-text.md · JSON https://www.anchorterminal.com/api/v1/tools/openai-speech-to-text.json - Best speech-to-text APIs for AI agents: https://www.anchorterminal.com/best/speech-to-text/index.md - All 91 stt comparisons: https://www.anchorterminal.com/compare/speech-to-text/index.md ## Which one, for what ### Cartesia Ink (B) Good for: Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech. Ahead on: - Payments & pricing, 35 against 20 - Maintenance & community, 85 against 75 - Transparency & trust, 67 against 61 Watch for: The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls ### OpenAI Speech to Text (BB) Good for: Suited to plain transcription of recorded files at a low price a minute and to teams already holding an OpenAI key. Ahead on: - Reliability, 80 against 73 - Security & auth, 86 against 52 Also in its favour: - Agent-ready, a grade of BB or better Watch for: `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe` and `gpt-4o-transcribe-diarize` were deprecated on 26 August 2026 and shut down on 26 February 2027 ## Score by category | Category | Weight | Cartesia Ink | OpenAI Speech to Text | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 73 | 80 | OpenAI Speech to Text +7 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 89 | 88 | Cartesia Ink +1 | | Agent ergonomics | 13% (16.2 this run) | 74 | 78 | OpenAI Speech to Text +4 | | Security & auth | 14% (17.5 this run) | 52 | 86 | OpenAI Speech to Text +34 | | Payments & pricing | 10% (12.5 this run) | 35 | 20 | Cartesia Ink +15 | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 85 | 75 | Cartesia Ink +10 | | Transparency & trust | 7% (8.8 this run) | 67 | 61 | Cartesia Ink +6 | | Negative events | ≤15 | 0 | 0 | | | **Total** | | **67.9 · B** | **72.4 · BB** | | ## Facts side by side | Fact | Cartesia Ink | OpenAI Speech to Text | | --- | --- | --- | | Kind | Model API | Model API | | Vendor | Cartesia | OpenAI | | Hosted endpoint | `https://api.cartesia.ai` | `https://api.openai.com/v1` | | Transports | HTTP, websocket, Streamable HTTP | HTTP, websocket | | Auth | API key | API key | | Pricing | Freemium | Pay per use | | x402 | no | no | | Licence | Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0 | Proprietary hosted service. The service terms were not read (see open questions). The Python SDK is Apache-2.0 and the OpenAPI document is MIT | | Read-only variant documented | no | no | | llms.txt | yes | yes | | Last release | 2026-09-17 | 2026-08-26 | | Terms last updated | 2024-06-14 | couldn't be read | | Privacy policy last updated | 2024-06-14 | couldn't be read | | Customer content may train models | yes, with an opt-out | couldn't be read | | Terms restrict automated access | yes | couldn't be read | | Terms restrict benchmarking | yes | couldn't be read | | Terms or service can change without notice | not found in the text | couldn't be read | | Arbitration or class-action waiver | yes | couldn't be read | | Popularity | 134 stars | 32k stars | ## Verdicts **Cartesia Ink.** Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation. **OpenAI Speech to Text.** `gpt-transcribe` costs $0.0045 an audio minute, and the audio endpoints keep no abuse-monitoring logs or application state. Speaker labels, timestamps, subtitles and translation exist only on `whisper-1` and `gpt-4o-transcribe-diarize`, which shut down on 26 February 2027 with no named replacement for those functions. ## Before you call either ### Cartesia Ink 1. Use `wss://api.cartesia.ai/stt/turns/websocket` with `model=ink-2`, `encoding`, `sample_rate` and `cartesia_version=2026-08-14`. All four are required. 2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error. 3. Read the final text from `turn.end` only. `transcript` is cumulative within a turn, so joining `turn.update` events duplicates text. 4. Send `{"type": "close"}` after the last audio and keep reading until the server closes the socket, or the buffered tail is lost. 5. Check `encoding` and `sample_rate` against the source before sending. The docs say the server might not return an error when they are wrong. ### OpenAI Speech to Text 1. Send `gpt-transcribe` to `POST /v1/audio/transcriptions` for recorded files. Use `languages` (a list), not `language`, and never send both. 2. Keep each upload at 25 MB or less. Split longer audio between sentences and pass the previous chunk's text in `prompt`. 3. For speaker labels send `gpt-4o-transcribe-diarize` with `response_format=diarized_json` and `chunking_strategy=auto` for audio over 30 seconds. Plan for its shutdown on 26 February 2027. 4. Word timestamps, `srt`, `vtt` and `/v1/audio/translations` need `whisper-1`, which cannot stream and shuts down on the same date. 5. On 429 or 503 wait at least `Retry-After` when present, then back off with jitter. Do not retry `credit_balance_exhausted` or spend-limit errors. ## Questions ### Which is better for AI agents, Cartesia Ink or OpenAI Speech to Text? OpenAI Speech to Text scores 72.4 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 3 of 7 scored categories. Cartesia Ink leads on payments & pricing, maintenance & community and transparency & trust. ### Do Cartesia Ink and OpenAI Speech to Text need an API key? Both need an API key. ### Can an agent call Cartesia Ink and OpenAI Speech to Text without installing anything? Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and OpenAI Speech to Text at https://api.openai.com/v1. ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.json, and with the fewest tokens: https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-openai-speech-to-text.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "cartesia-ink-stt", "b": "openai-speech-to-text"}`. From a terminal: `anchor compare cartesia-ink-stt openai-speech-to-text` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json and https://www.anchorterminal.com/api/v1/tools/openai-speech-to-text.json ## Other comparisons with Cartesia Ink or OpenAI Speech to Text - [Amazon Transcribe vs Cartesia Ink](https://www.anchorterminal.com/compare/amazon-transcribe-vs-cartesia-ink-stt.md) - [Amazon Transcribe vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/amazon-transcribe-vs-openai-speech-to-text.md) - [AssemblyAI Speech-to-Text (Universal) vs Cartesia Ink](https://www.anchorterminal.com/compare/assemblyai-stt-vs-cartesia-ink-stt.md) - [AssemblyAI Speech-to-Text (Universal) vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/assemblyai-stt-vs-openai-speech-to-text.md) - [Azure AI Speech speech-to-text vs Cartesia Ink](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-cartesia-ink-stt.md) - [Azure AI Speech speech-to-text vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/azure-speech-to-text-vs-openai-speech-to-text.md) - [Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux)](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-deepgram-stt.md) - [Cartesia Ink vs ElevenLabs Scribe Speech to Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-elevenlabs-scribe.md) - [Cartesia Ink vs Gladia Speech-to-Text API + MCP](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-gladia-stt.md) - [Cartesia Ink vs Google Cloud Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-google-speech-to-text.md) - [Cartesia Ink vs Groq Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-groq-speech-to-text.md) - [Cartesia Ink vs Mistral Voxtral Transcribe](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-mistral-voxtral-transcribe.md) - [Cartesia Ink vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-rev-ai-stt.md) - [Cartesia Ink vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-soniox-stt.md) - [Cartesia Ink vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/cartesia-ink-stt-vs-speechmatics-stt.md) - [Deepgram Speech-to-Text (Nova-3, Flux) vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/deepgram-stt-vs-openai-speech-to-text.md) - [ElevenLabs Scribe Speech to Text API vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/elevenlabs-scribe-vs-openai-speech-to-text.md) - [Gladia Speech-to-Text API + MCP vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/gladia-stt-vs-openai-speech-to-text.md) - [Google Cloud Speech-to-Text vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/google-speech-to-text-vs-openai-speech-to-text.md) - [Groq Speech-to-Text vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/groq-speech-to-text-vs-openai-speech-to-text.md) - [Mistral Voxtral Transcribe vs OpenAI Speech to Text](https://www.anchorterminal.com/compare/mistral-voxtral-transcribe-vs-openai-speech-to-text.md) - [OpenAI Speech to Text vs Rev AI Speech-to-Text API](https://www.anchorterminal.com/compare/openai-speech-to-text-vs-rev-ai-stt.md) - [OpenAI Speech to Text vs Soniox Speech-to-Text](https://www.anchorterminal.com/compare/openai-speech-to-text-vs-soniox-stt.md) - [OpenAI Speech to Text vs Speechmatics Speech-to-Text](https://www.anchorterminal.com/compare/openai-speech-to-text-vs-speechmatics-stt.md)