# Cartesia Ink (slim) > Cartesia's hosted speech-to-text API. Ink 2 transcribes live audio in five languages over a WebSocket with built-in turn detection, and the older Ink Whisper model transcribes uploaded files in about 100 languages. - Full: https://www.anchorterminal.com/tools/cartesia-ink-stt.md (~9,350 tokens) · this version ~1,730 tokens · JSON https://www.anchorterminal.com/tools/cartesia-ink-stt.json · canonical https://www.anchorterminal.com/tools/cartesia-ink-stt - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **B · 67.9/100 · rank #238 of 950 · #9 in Speech-to-text · not agent-ready · confidence medium** Assessment: Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation. ## Facts - Kind: Model API · vendor: Cartesia · category: Speech-to-text · legal entity: Cartesia AI, Inc. · provenance 84/100 - Endpoint: `https://api.cartesia.ai` (HTTP, websocket, Streamable HTTP) - Auth: API key · pricing: Freemium · x402: no · licence: Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - Models: `ink-2` (Stable) and `ink-preview` (Preview) for realtime, `ink-whisper` for realtime and batch - Languages: `ink-2` covers English, French, Hindi, Japanese and Spanish with automatic detection. `ink-whisper` lists about 100 language codes - Streaming latency: Cartesia claims a 0.1 second time to final transcript on its product page. Not measured by us - Turn detection: Built into `ink-2` on `/stt/turns/websocket`, with start, eager-end and end thresholds and an end timeout of 640 to 11,200 ms (default 5,600) - Diarisation: None found - Input: Realtime takes raw mono PCM (`pcm_s16le`, `pcm_s32le`, `pcm_f16le`, `pcm_f32le`, `pcm_mulaw`, `pcm_alaw`). Batch takes flac, m4a, mp3, mp4, mpeg, mpga, oga, ogg, wav or webm of any length - Output: Turn events with cumulative transcripts on the auto endpoint, transcript deltas with `is_final` on the manual endpoint, one transcript in batch. Word timestamps on the manual and batch endpoints - Free tier: 20,000 credits a month, about 1 hour 51 minutes of `ink-2` audio, for non-commercial use - Rate limits: Concurrent STT requests 8 (Free), 12 (Pro), 20 (Startup), 60 (Scale), custom on Enterprise. Idle sockets close after 3 minutes - Trains on API data: Yes unless otherwise agreed, per the terms. An opt-out form is in the Playground's data controls - Data retention: Zero data retention is an Enterprise setting that covers STT audio and transcripts. No period was found for other plans - Compliance: The docs and pricing page state SOC 2 Type II, HIPAA, PCI-DSS service provider and GDPR. The trust centre could not be read - Regions: `api.cartesia.ai` routes to the closest endpoint. Dedicated US, EU, UK, India and Australia deployments for Enterprise - MCP: Hosted at `https://mcp.cartesia.ai/mcp` with a Playground sign-in, or run locally. 18 tools in `server.py` at v0.26.1, one of them `speech_to_text` - Prices: Ink 2 realtime, Scale plan $0.0067 per minute of audio; Ink 2 realtime, Startup plan $0.0071 per minute of audio; Ink 2 realtime, Pro plan $0.009 per minute of audio; Ink Whisper realtime, Scale plan $0.0022 per minute of audio; Ink Whisper batch, Scale plan $0.0011 per minute of audio - Scores: Reliability 73, Performance pending, Schema & documentation 89, Agent ergonomics 74, Security & auth 52, Payments & pricing 35, Task success pending, Maintenance & community 85, Transparency & trust 67 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading. · Agent ergonomics, API reading, as for the other speech-to-text listings. · Security & auth, Model reading. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Model reading, as for the other hosted speech-to-text listings. · Transparency & trust, The service is closed and the SDKs are Apache-2.0. The Terms of Service cover the website, its subdomains and APIs, were last revised on 14… - Sources: 36, open questions: 15, both in the full twin - Capabilities: speech.stt, speech.streaming, speech.batch, speech.languages - JSON: https://www.anchorterminal.com/api/v1/tools/cartesia-ink-stt.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/cartesia-ink-stt.svg` or a link to https://www.anchorterminal.com/tools/cartesia-ink-stt from a page on cartesia.ai or one of its subdomains, or the README of github.com/cartesia-ai/cartesia-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Use `wss://api.cartesia.ai/stt/turns/websocket` with `model=ink-2`, `encoding`, `sample_rate` and `cartesia_version=2026-08-14`. All four are required. 2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error. 3. Read the final text from `turn.end` only. `transcript` is cumulative within a turn, so joining `turn.update` events duplicates text. 4. Send `{"type": "close"}` after the last audio and keep reading until the server closes the socket, or the buffered tail is lost. 5. Check `encoding` and `sample_rate` against the source before sending. The docs say the server might not return an error when they are wrong. ## Connect ```bash pip install 'cartesia[websockets]' # or: npm install @cartesia/cartesia-js ``` ```bash claude mcp add --transport http --scope user cartesia https://mcp.cartesia.ai/mcp ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Amazon Transcribe | BB | 73.4 | speech.stt, speech.streaming, speech.batch, speech.languages | https://www.anchorterminal.com/tools/amazon-transcribe.min.md | | Azure AI Speech speech-to-text | BB | 73 | speech.stt, speech.streaming, speech.batch, speech.languages | https://www.anchorterminal.com/tools/azure-speech-to-text.min.md | | OpenAI Speech to Text | BB | 72.4 | speech.stt, speech.streaming, speech.batch, speech.languages | https://www.anchorterminal.com/tools/openai-speech-to-text.min.md | | Deepgram Speech-to-Text (Nova-3, Flux) | BB | 70.3 | speech.stt, speech.streaming, speech.batch, speech.languages | https://www.anchorterminal.com/tools/deepgram-stt.min.md | | Google Cloud Speech-to-Text | BB | 70.2 | speech.stt, speech.streaming, speech.batch, speech.languages | https://www.anchorterminal.com/tools/google-speech-to-text.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)