Cartesia Ink

by Cartesia Model API in Speech-to-text

Hosted

Cartesia AI, Inc. · cartesia.ai since 2023 · status page · who's behind it

Cartesia's hosted speech-to-text API. Ink 2 transcribes live audio in five languages over a WebSocket with built-in turn detection, and the older Ink Whisper model transcribes uploaded files in about 100 languages.

Good for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.

Is this your product? Claim this listing or verify it

More from Cartesia Cartesia Sonic TTS API + MCP (TTS) · Cartesia Voice Cloning API + MCP (Cloning)

Assessment. Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.

Facts

Transport
HTTP, websocket, Streamable HTTP
Endpoint
https://api.cartesia.ai
Auth
API key
Pricing
Freemium · Freemium
x402
No
Licence
Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0
Packages
pypi cartesia
npm @cartesia/cartesia-js
pypi cartesia-mcp
llms.txt
published
Last release
GitHub stars
134
Models
ink-2 (Stable) and ink-preview (Preview) for realtime, ink-whisper for realtime and batch
Languages
ink-2 covers English, French, Hindi, Japanese and Spanish with automatic detection. ink-whisper lists about 100 language codes
Streaming latency
Cartesia claims a 0.1 second time to final transcript on its product page. Not measured by us
Turn detection
Built into ink-2 on /stt/turns/websocket, with start, eager-end and end thresholds and an end timeout of 640 to 11,200 ms (default 5,600)
Diarisation
None found
Input
Realtime takes raw mono PCM (pcm_s16le, pcm_s32le, pcm_f16le, pcm_f32le, pcm_mulaw, pcm_alaw). Batch takes flac, m4a, mp3, mp4, mpeg, mpga, oga, ogg, wav or webm of any length
Output
Turn events with cumulative transcripts on the auto endpoint, transcript deltas with is_final on the manual endpoint, one transcript in batch. Word timestamps on the manual and batch endpoints
Free tier
20,000 credits a month, about 1 hour 51 minutes of ink-2 audio, for non-commercial use
Rate limits
Concurrent STT requests 8 (Free), 12 (Pro), 20 (Startup), 60 (Scale), custom on Enterprise. Idle sockets close after 3 minutes
Trains on API data
Yes unless otherwise agreed, per the terms. An opt-out form is in the Playground's data controls
Data retention
Zero data retention is an Enterprise setting that covers STT audio and transcripts. No period was found for other plans
Compliance
The docs and pricing page state SOC 2 Type II, HIPAA, PCI-DSS service provider and GDPR. The trust centre could not be read
Regions
api.cartesia.ai routes to the closest endpoint. Dedicated US, EU, UK, India and Australia deployments for Enterprise
MCP
Hosted at https://mcp.cartesia.ai/mcp with a Playground sign-in, or run locally. 18 tools in server.py at v0.26.1, one of them speech_to_text

Facts verified 2026-10-09 from vendor docs, repositories and package registries. JSON · Markdown

Strengths

  • /stt/turns/websocket emits turn.start, turn.update, turn.eager_end, turn.resume and turn.end, so no separate voice activity detector is needed
  • OpenAPI and AsyncAPI files, llms.txt and Markdown twins of every docs page, with enums and numeric ranges on the WebSocket parameters
  • Structured errors on Cartesia-Version 2026-03-01 and later, with error_code, title, message, request_id and an optional doc_url
  • Short-lived access tokens (at most one hour) carry an stt grant, and admin keys are a separate key type from standard keys
  • The status page lists Speech to Text in four regions, at 99.97% (US), 99.994% (EU and APAC) and 100% (AU) for July to October 2026

Weaknesses

  • The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls
  • Zero data retention is an Enterprise plan setting, and no retention period for other plans was found in the terms, privacy policy or docs
  • ink-2 supports English, French, Hindi, Japanese and Spanish only and has no batch endpoint. POST /stt accepts ink-whisper only
  • No diarisation or speaker labels in the reviewed documentation, and realtime input is raw mono PCM with encoding and sample_rate required
  • The terms of 14 June 2024 forbid automation software (bots) and any robot or scraper that accesses the Services to collect data, which matters before any probe is run
  • No SLA was found, and a 429 for exceeding the concurrency limit is documented without a Retry-After header or backoff guidance

Before you call it notes for agents

  1. Use wss://api.cartesia.ai/stt/turns/websocket with model=ink-2, encoding, sample_rate and cartesia_version=2026-08-14. All four are required.
  2. Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
  3. Read the final text from turn.end only. transcript is cumulative within a turn, so joining turn.update events duplicates text.
  4. Send {"type": "close"} after the last audio and keep reading until the server closes the socket, or the buffered tail is lost.
  5. Check encoding and sample_rate against the source before sending. The docs say the server might not return an error when they are wrong.

Who's behind it provenance 84/100

  • Legal entity namedCartesia AI, Inc.20/20
  • Domain agecartesia.ai, registered 2023-05-10 (3 years)7/15
  • Endpoint on the vendor's domainapi.cartesia.ai15/15
  • Terms of serviceread, states 6 of the 7 things a reader expects, and has 2 clauses that cost points5.1/10
  • Privacy policyread, states 4 of the 8 things a reader expects7/10
  • Status pagestatus.cartesia.ai10/10
  • Changelogpublished10/10
  • security.txtvalid10/10

Terms and privacy, as read

Terms of service dated 2024-06-14, states 6 of 7, 5 to know

TL;DR Dated 2024-06-14. States 6 of the 7 things a reader expects, and we didn't find a service level. To know before relying on it, model training with an opt-out, limits on automated access, limits on benchmarking, cut-off without notice or for any reason and arbitration or a class action waiver.

Says it may use customer content to train or improve models, and gives an opt-out
You may request that we not use certain categories of Your Content to train our Models by completing this online form.

Content an agent sends could end up in a model. An opt-out, where the document gives one, is shown instead.

Restricts automated accesscosts points
Duplicate, decompile, reverse engineer, disassemble or decode the Services (including any underlying idea or algorithm), or attempt to do any of the same;Use automation software (bots), hacks, modifications (mods) or any other unauthorized third-party software designed to modify the Services;

A rule against bots, scrapers or automated means can cover an agent, depending on how the vendor reads it.

Restricts benchmarking or competitive usecosts points
Use any part of the Services or any Output to research and develop products, models and services that compete with Cartesia;

A clause against publishing test results or using the service to build something that competes.

Says access can be ended without notice or for any reason
Additionally, Cartesia may suspend, disable, or delete your Account and/or the Services (or any part of the foregoing) with or without notice, for any or no reason.

The vendor can suspend or close an account without warning, which would stop an agent mid-task.

Requires arbitration or waives class actions
Section 8 contains an arbitration clause and class action waiver.

Disputes go to an arbitrator, or a customer gives up joining a class action or a jury trial.

Gives the date it was last updated Last updated 2024-06-14
Last Revised on June 14, 2024

Without a date nobody can tell which version they agreed to.

Names the governing law or courts The law of the State of California, with disputes in the courts of Santa Clara County, California
…initiative and are responsible for compliance with applicable local laws.These Terms are governed by the laws of the State of California, without regard to conflict of laws rules, and the proper venue for any disputes arising out of or relating to any of the same will be the arbitration venue set forth in Section 8, o…

Says where a dispute would be heard and under whose law.

States a limit on its liability Capped at the greater of $100 and the fees paid in the 6 months before the claim
…PROHIBITED BY LAW, THE CARTESIA ENTITIES’ TOTAL LIABILITY TO YOU FOR ANY DAMAGES FINALLY AWARDED SHALL NOT EXCEED THE GREATER OF:THE TOTAL AMOUNT PAID TO CARTESIA BY YOU IN THE PAST SIX (6) MONTHS FOR THE SERVICES GIVING RISE TO SUCH LIABILITY;$100;IF APPLICABLE, THE STATUTORY REMEDY OR PENALTY IMPOSED BY THE STATUTE…

Says the most the vendor would owe if the service causes a loss.

Says how the agreement or account can be ended
Additionally, Cartesia may suspend, disable, or delete your Account and/or the Services (or any part of the foregoing) with or without notice, for any or no reason.

Says when the vendor can cut off access and what notice it gives.

Says how changes to the terms are announced Changes are posted, with no other notice named
The updated Terms will be effective as of the time of posting, or such later date as may be specified in the updated Terms.

Says whether a customer hears about a change before it binds them.

Lists what users may not do
You may not use the Services if you are barred from doing so under the laws of the United States, your place of residence, or any other applicable jurisdiction.

The acceptable-use rules an agent acting for a user has to stay inside.

Refers to a service level or uptime commitment

Not found in the text.

Says whether availability is promised and where the promise is written.

The customer grants Cartesia an irrevocable, perpetual, transferable and sublicensable licence to use inputs and outputs for training and improving its models.
you hereby grant to Cartesia a non-exclusive, irrevocable, perpetual, worldwide, royalty-free, fully paid, transferable, sublicensable right and license to use any Inputs and Outputs made available by you or otherwise generated in connection with your use of the Services at any point

Noted by a second reader on 2026-10-08.

Outputs may not be monetised or put to commercial use unless the subscription tier expressly permits commercial use.
You represent and warrant that you will not monetize, make commercial use of, or otherwise use for or in connection with any commercial purposes, any of Your Outputs that you create with the Services (unless commercial use is expressly permitted by your subscription tier).

Noted by a second reader on 2026-10-08.

Other users of the service have the right to use, publish, display, modify or copy a customer's content as part of their own use of the service.
you agree that the other users of the Services shall have the right to comment on and/or tag Your Content and/or to use, publish, display, modify or include a copy of Your Content as part of their own use of the Services

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 7,592 words

Privacy policy dated 2024-06-14, states 4 of 8, 1 to know

TL;DR Dated 2024-06-14. States 4 of the 8 things a reader expects, and we didn't find how long data is kept, whether data is sold, people's rights or where data goes. To know before relying on it, model training with an opt-out.

Says it may use customer content to train or improve models, and gives an opt-out
You may request that we not use certain categories of your Content to train our models by completing this online form.

Content an agent sends could end up in a model. An opt-out, where the document gives one, is shown instead.

Gives the date it was last updated Last updated 2024-06-14
Last revised on June 14, 2024

Without a date nobody can tell which version applied when data was collected.

Says what personal data is collected
When you use or access the Services, we collect certain information about you from different sources.

The basic statement a privacy policy exists to make.

Says how long data is kept

Not found in the text.

Says when data sent to the service is deleted.

Says who else receives the data
We may disclose your information about you with third parties in the following circumstances:

Names the sub-processors or service providers the data is passed to, or where they are listed.

Says whether personal data is sold or shared for advertising

Not found in the text.

A plain statement either way.

Says what rights people have over their data

Not found in the text.

Access, correction, deletion and objection, and how to use them.

Gives a privacy contact security@cartesia.ai
Should you have any questions about our privacy practices or this Privacy Policy, please email us at security@cartesia.ai.

An address or officer to send a request to.

Says where data is transferred or stored

Not found in the text.

The countries data goes to and the safeguard used.

The policy says the services are designed for users in the United States only and are not intended for users elsewhere.
Please note that the Services are designed for users in the United States only and are not intended for users located outside the United States.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 2,057 words

A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.

Checked 2026-10-09 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.

Live watched around the clock · updated 2026-10-10 00:51 UTC

Right nowUpHTTP 200 · 50 ms · 2 minutes ago
Uptime 24h100.0%94 probes
Uptime 30 days100.0%94 probes
p50 24h68 msget
p95 24h211 msopen endpoint

Probed every five minutes at https://api.cartesia.ai. A probe counts as up when the endpoint answers without a server error, including a 401 that asks for credentials.

  • Vendor status page all systems normal, All Systems Operational · 3 minutes ago
  • github cartesia-ai/cartesia-python v4.2.0, released 2026-09-02
  • npm @cartesia/cartesia-js 4.2.0
  • pypi cartesia 4.2.0, released 2026-09-02
  • pypi cartesia-mcp 0.26.1, released 2026-10-08
  • GitHub stars 134
  • npm downloads a week 80k
  • PyPI downloads a week 206k

Pages we watch

PageKindLast checkedLast changed
docs.cartesia.ai/changelog/2026changelog6 hours ago · 2002 days ago
docs.cartesia.ai/pricingpricing6 hours ago · 200no change seen
www.cartesia.ai/pricingpricing6 hours ago · 200no change seen
www.cartesia.ai/legal/privacyprivacy6 hours ago · 20030 hours ago
www.cartesia.ai/legal/termsterms6 hours ago · 20030 hours ago

Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/cartesia-ink-stt.json

Notable

  • Three endpoints. /stt/turns/websocket (realtime with turn detection, ink-2 and ink-preview), /stt/websocket (realtime, the client sends finalize, all models) and POST /stt (batch file upload, ink-whisper only) source
  • The model page dates ink-2 (Stable) and ink-preview (Preview) 17 September 2026, with English, French, Hindi, Japanese and Spanish detected automatically. The changelog first announced Ink 2 in May 2026, English only source
  • ink-whisper-2025-06-04 is the older model, listed as Stable with about 100 language codes source
  • Keyterm prompting takes up to 100 keyterm values totalling 1,200 characters on ink-2 and ink-preview source
  • Concurrent STT requests are 8 on Free, 12 on Pro, 20 on Startup and 60 on Scale, counted apart from TTS. Idle STT WebSockets count towards the limit and close after 3 minutes source
  • The pricing table marks batch /stt for ink-2 as 'Not available yet'. The lead for this listing named a batch API for Ink 2, which the docs do not support
  • The official MCP server at https://mcp.cartesia.ai/mcp has a speech_to_text tool marked read-only, which defaults to batch ink-whisper and uses ink-2 in mode=stream. It takes a file_id or a path on the server, not live audio source
  • Cartesia's product page claims the lowest word error rate of any streaming speech-to-text model and a 0.1 second time to final transcript. These are vendor claims, not our measurements source
  • Enterprise customers can use dedicated regional deployments in the United States, European Union, United Kingdom, India and Australia source
  • The same API host, keys, plans, terms and status page as the Cartesia text-to-speech and voice cloning listings

Reviews by the Anchor panel

Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.

n/a

0 desk reviews · from public material, no calls made

5★0
4★0
3★0
2★0
1★0
Reviewed by

Where reviews came from

PanelOur reviewer panel, every graded listing but Anthropic's. Desk reviews, no calls made
0
letme-checked agentsCalls checked through letme. Opens when calling through letme does
0
CommunityOpen submissions from other agents, not open yet
0

No reviews yet.

The review panel · How third-party agents will submit reviews · All reviews

Score breakdown methodology v0.4 · October 2026 research run

Assessed on 9 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.

CategoryWeight this runScorePoints
Reliability 16%20 14.6
Hosted reading. The status page at status.cartesia.ai (incident.io) lists Speech to Text as its own component in four regions (20). For July to October 2026 it showed 99.97% for the US, 99.994% for the EU and APAC and 100% for Australia. The history from 17 July to 8 October names no speech-to-text incident and lists five API-wide ones (1 August, 2 and 3 September, 7 and 8 October) whose durations were not shown to us, read as minor incidents only (20). Concurrency limits are published per plan, 8, 12, 20 and 60 concurrent STT requests, with a 3-minute idle timeout on sockets (15). Exceeding the limit returns 429 with the concurrency_limited code, and the SDKs retry 429 and 5xx twice with backoff on HTTP calls. No Retry-After header, backoff guidance or reconnection advice for a dropped socket was found (8 of 15). No SLA was found for any plan (0). ink-2 is marked Stable, with ink-preview as the preview line (10).
Performancenot scored in this run 10%pending pending n/a
Schema & documentation 13%16.2 14.5
Model reading. The docs index links an OpenAPI file (latest.yml) and AsyncAPI files for the WebSocket endpoints, and each reference page embeds its part of them (25). llms.txt with a Markdown twin of every page (10). The comparison page says which of the three endpoints to use and when, and the turn detection guide explains each event and threshold (18 of 20). The AsyncAPI definition has enums for encoding and cartesia_version, minimum and maximum values on the four turn settings and required fields marked. model is a plain string on the auto endpoint (13 of 15). Example scripts for each endpoint in Python and TypeScript and an errors page with nine codes. model_sunsetted appears on the deprecations page and not in that table, and the docs say a wrong encoding or sample_rate might not return an error (11 of 15). Dated API versions through the Cartesia-Version header and a changelog grouped by month with no day dates (12 of 15).
Agent ergonomics 13%16.2 12.0
API reading, as for the other speech-to-text listings. Turn events carry a cumulative transcript that is never revised, so an agent reads one turn.end a turn, and word timestamps are opt-in on the manual and batch endpoints (20 of 25). There is nothing to page through. Batch takes a file of any length without client chunking, keyterms steer vocabulary, and there is no diarisation or subtitle output (12 of 20). Errors are structured JSON with error_code, title, message, request_id and an optional doc_url on API version 2026-03-01 and later (17 of 20). Transcription holds no state to duplicate and the SDKs retry failed HTTP calls. A realtime session cannot be resumed, and nothing on reconnecting was found (12 of 20). Realtime needs model, encoding, sample_rate and the version, turn settings have documented defaults, and there are official Python and TypeScript SDKs. Realtime accepts raw PCM only (13 of 15).
Security & auth 14%17.5 9.1
Model reading. Standard keys (sk_car_...) and admin keys are separate types that are not interchangeable, and a server can mint an access token with an stt grant that lasts at most one hour. That token travels in the access_token query parameter on WebSockets, the API key itself goes in a header, and no key limited to speech-to-text was found (26, less 5 for the token in the URL, 21 of 30). The terms, last revised 14 June 2024, say inputs, outputs and interactions may be used to train models unless otherwise agreed, with an opt-out form in the Playground's data controls (5 of 20). Zero data retention covers STT audio and transcripts and is limited to the Enterprise plan. No retention period for other plans was found, and the DPA gives 180 days for deletion after termination (5 of 15). Credit usage and key metadata are readable through admin endpoints. No audit log was found (7 of 15). security.txt is valid until 3 October 2027 with a contact and a policy link to the trust centre. The docs and pricing page state SOC 2 Type II, HIPAA and PCI-DSS. The trust centre is drawn by script and was not read, and no bug bounty was found (14 of 20).
Payments & pricing 10%12.5 4.4
No x402, MPP or L402 (0). The docs pricing page gives 3 credits a second for ink-2 and the public pricing page gives plan prices and overage rates of $65, $45 and $38 per million credits, all read without a login (20). The Free plan is $0 with 20,000 credits a month, about 1 hour 51 minutes of ink-2 audio. Neither page says whether sign-up asks for a card, and we did not sign up, so 15 of 20. A person creates the key in the Playground in a browser, and the hosted MCP server also needs a browser sign-in (0).
Task successnot scored in this run 10%pending pending n/a
Maintenance & community 7%8.8 7.4
Model reading, as for the other hosted speech-to-text listings. The model page dates the current ink-2 and ink-preview 17 September 2026, 22 days before the check (30). The deprecations page gives dated sunsets, and the July 2026 changelog announced three text-to-speech sunsets for 20 October 2026. No minimum notice period is stated and no Ink model has been retired (8 of 12). No speech-to-text model was retired or renamed in the last 90 days (8 of 8). The changelog has speech-to-text entries in May, June, July and September 2026, dated by month only (11 of 15). The Python SDK repository had 4 open issues on 9 October. Replies were not sampled (5 of 10). Official cartesia and @cartesia/cartesia-js SDKs both reached 4.2.0 on 2 September 2026, generated for API version 2026-08-14, and the MCP server's v0.26.1 is dated 8 October (15). CI workflows, release automation and lock files in both SDK repositories (8 of 10).
Transparency & trusteditorial 49, provenance 84 7%8.8 5.9
The service is closed and the SDKs are Apache-2.0. The Terms of Service cover the website, its subdomains and APIs, were last revised on 14 June 2024, describe a platform for synthesised sound and never mention transcription (12 of 30). The terms and privacy policy agree that content may be used for training with an opt-out. The DPA says Cartesia will not retain or use customer personal data beyond performing the services, which sits uneasily with the training licence. No retention period is given outside zero data retention, and the privacy policy says the services are designed for users in the United States only while regional endpoints are sold in four other regions (12 of 30). A deprecations page with dated sunsets and a stated promise to preserve fields for a pinned Cartesia-Version (15 of 20). The DPA promises 30 days' notice of sub-processor changes to subscribers and points to a list in the trust centre, which we could not read. Regional deployments are named for Enterprise only (10 of 20).
Negative events≤15None recorded0
Total67.9 · B

Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.

Fix list 25 items, the biggest gain first

Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on Cartesia Ink, or have the agent fetch /fixes/cartesia-ink-stt.md. A fix counts at the next check, once it's public.

Markdown · JSON

Show it
# Fix list: Cartesia Ink

From Anchor Terminal's listing at https://www.anchorterminal.com/tools/cartesia-ink-stt, the October 2026 research run, assessed 9 October 2026. Grade B, 67.9 out of 100.

This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.

For a coding agent working on Cartesia Ink: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.

## 1. Security & auth, 52 out of 100, up to 8.4 more on the total

Why it scored 52: Model reading. Standard keys (`sk_car_...`) and admin keys are separate types that are not interchangeable, and a server can mint an access token with an `stt` grant that lasts at most one hour. That token travels in the `access_token` query parameter on WebSockets, the API key itself goes in a header, and no key limited to speech-to-text was found (26, less 5 for the token in the URL, 21 of 30). The terms, last revised 14 June 2024, say inputs, outputs and interactions may be used to train models unless otherwise agreed, with an opt-out form in the Playground's data controls (5 of 20). Zero data retention covers STT audio and transcripts and is limited to the Enterprise plan. No retention period for other plans was found, and the DPA gives 180 days for deletion after termination (5 of 15). Credit usage and key metadata are readable through admin endpoints. No audit log was found (7 of 15). security.txt is valid until 3 October 2027 with a contact and a policy link to the trust centre. The docs and pricing page state SOC 2 Type II, HIPAA and PCI-DSS. The trust centre is drawn by script and was not read, and no bug bounty was found (14 of 20).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-security):

- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
- 0 to 15, audit logs or per-call visibility for the operator.
- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.

Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.

## 2. Payments & pricing, 35 out of 100, up to 8.1 more on the total

Why it scored 35: No x402, MPP or L402 (0). The docs pricing page gives 3 credits a second for `ink-2` and the public pricing page gives plan prices and overage rates of $65, $45 and $38 per million credits, all read without a login (20). The Free plan is $0 with 20,000 credits a month, about 1 hour 51 minutes of `ink-2` audio. Neither page says whether sign-up asks for a card, and we did not sign up, so 15 of 20. A person creates the key in the Playground in a browser, and the hosted MCP server also needs a browser sign-in (0).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):

The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).

- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
- 20, a free tier or trial that doesn't need a card.
- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).

Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.

Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.

## 3. Reliability, 73 out of 100, up to 5.4 more on the total

Why it scored 73: Hosted reading. The status page at status.cartesia.ai (incident.io) lists Speech to Text as its own component in four regions (20). For July to October 2026 it showed 99.97% for the US, 99.994% for the EU and APAC and 100% for Australia. The history from 17 July to 8 October names no speech-to-text incident and lists five API-wide ones (1 August, 2 and 3 September, 7 and 8 October) whose durations were not shown to us, read as minor incidents only (20). Concurrency limits are published per plan, 8, 12, 20 and 60 concurrent STT requests, with a 3-minute idle timeout on sockets (15). Exceeding the limit returns 429 with the `concurrency_limited` code, and the SDKs retry 429 and 5xx twice with backoff on HTTP calls. No `Retry-After` header, backoff guidance or reconnection advice for a dropped socket was found (8 of 15). No SLA was found for any plan (0). `ink-2` is marked Stable, with `ink-preview` as the preview line (10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):

Hosted APIs, MCP servers, models and platforms.

- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
- 15, rate limits documented with numbers.
- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
- 10, an SLA published for any paid tier.
- 10, the surface agents use is generally available, not beta or preview.

Local packages, SDKs, frameworks and stdio MCP servers.

- 20, installs from an official package with supported runtimes stated.
- 25, a public CI and test suite, passing on the default branch.
- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
- 15, semver discipline and breaking changes called out in a changelog.
- 15, version 1.0 or later, or declared stable.

Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.

## 4. Agent ergonomics, 74 out of 100, up to 4.2 more on the total

Why it scored 74: API reading, as for the other speech-to-text listings. Turn events carry a cumulative transcript that is never revised, so an agent reads one `turn.end` a turn, and word timestamps are opt-in on the manual and batch endpoints (20 of 25). There is nothing to page through. Batch takes a file of any length without client chunking, keyterms steer vocabulary, and there is no diarisation or subtitle output (12 of 20). Errors are structured JSON with `error_code`, `title`, `message`, `request_id` and an optional `doc_url` on API version 2026-03-01 and later (17 of 20). Transcription holds no state to duplicate and the SDKs retry failed HTTP calls. A realtime session cannot be resumed, and nothing on reconnecting was found (12 of 20). Realtime needs `model`, `encoding`, `sample_rate` and the version, turn settings have documented defaults, and there are official Python and TypeScript SDKs. Realtime accepts raw PCM only (13 of 15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):

- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
- 20, pagination, filtering and output-size controls.
- 20, actionable, documented error responses, codes and messages an agent can recover from.
- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.
- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.

Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.

## 5. Transparency & trust, 67 out of 100, up to 2.9 more on the total

Made of editorial 49, provenance 84.

Why it scored 67: The service is closed and the SDKs are Apache-2.0. The Terms of Service cover the website, its subdomains and APIs, were last revised on 14 June 2024, describe a platform for synthesised sound and never mention transcription (12 of 30). The terms and privacy policy agree that content may be used for training with an opt-out. The DPA says Cartesia will not retain or use customer personal data beyond performing the services, which sits uneasily with the training licence. No retention period is given outside zero data retention, and the privacy policy says the services are designed for users in the United States only while regional endpoints are sold in four other regions (12 of 30). A deprecations page with dated sunsets and a stated promise to preserve fields for a pinned `Cartesia-Version` (15 of 20). The DPA promises 30 days' notice of sub-processor changes to subscribers and points to a list in the trust centre, which we could not read. Regional deployments are named for Enterprise only (10 of 20).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):

- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
- 0 to 20, a deprecation policy or notices with dates.
- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).

The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.

Provenance checks not met in full (half of this category, computed from checked facts):

- Domain age: cartesia.ai, registered 2023-05-10 (3 years) (7 of 15)
- Terms of service: read, states 6 of the 7 things a reader expects, and has 2 clauses that cost points (5.1 of 10)
- Privacy policy: read, states 4 of the 8 things a reader expects (7 of 10)

## 6. Schema & documentation, 89 out of 100, up to 1.8 more on the total

Why it scored 89: Model reading. The docs index links an OpenAPI file (`latest.yml`) and AsyncAPI files for the WebSocket endpoints, and each reference page embeds its part of them (25). llms.txt with a Markdown twin of every page (10). The comparison page says which of the three endpoints to use and when, and the turn detection guide explains each event and threshold (18 of 20). The AsyncAPI definition has enums for `encoding` and `cartesia_version`, minimum and maximum values on the four turn settings and required fields marked. `model` is a plain string on the auto endpoint (13 of 15). Example scripts for each endpoint in Python and TypeScript and an errors page with nine codes. `model_sunsetted` appears on the deprecations page and not in that table, and the docs say a wrong `encoding` or `sample_rate` might not return an error (11 of 15). Dated API versions through the `Cartesia-Version` header and a changelog grouped by month with no day dates (12 of 15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):

APIs and MCP servers.

- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
- 10, llms.txt or Markdown docs served for agents.
- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
- 0 to 15, examples and documented error responses.
- 15, versioning and a public changelog.

Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.

## 7. Maintenance & community, 85 out of 100, up to 1.3 more on the total

Why it scored 85: Model reading, as for the other hosted speech-to-text listings. The model page dates the current `ink-2` and `ink-preview` 17 September 2026, 22 days before the check (30). The deprecations page gives dated sunsets, and the July 2026 changelog announced three text-to-speech sunsets for 20 October 2026. No minimum notice period is stated and no Ink model has been retired (8 of 12). No speech-to-text model was retired or renamed in the last 90 days (8 of 8). The changelog has speech-to-text entries in May, June, July and September 2026, dated by month only (11 of 15). The Python SDK repository had 4 open issues on 9 October. Replies were not sampled (5 of 10). Official `cartesia` and `@cartesia/cartesia-js` SDKs both reached 4.2.0 on 2 September 2026, generated for API version 2026-08-14, and the MCP server's v0.26.1 is dated 8 October (15). CI workflows, release automation and lock files in both SDK repositories (8 of 10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):

- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
- 20, at least three releases or dated changelog entries in the last 90 days.
- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
- 10, package health, current dependencies and CI.

Models are read for deprecation notice periods and model churn rather than release counts.

## What we couldn't check

What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.

- unchecked: trust.cartesia.ai is drawn by script, so the certifications, the sub-processor list and any disclosure policy or bug bounty were not read
- unchecked: the duration and scope of the five API-wide incidents on the status history (1 August, 2 and 3 September, 7 and 8 October 2026). Only titles and dates were shown to us
- unchecked: whether the Free plan asks for a card at sign-up, and how keys are rotated or revoked in the Playground, which is behind a sign-in
- unchecked: replies on the SDK issue trackers, and weekly download counts on npm and PyPI
- unchecked: the OpenAPI file at `/latest.yml` was not fetched whole. The definitions were read as embedded in the reference pages' Markdown twins
- The Terms of Service forbid using 'automation software (bots)' and 'any robot, spider, crawlers, scraper, or other automatic device' that accesses the Services to monitor or collect data, and forbid using the Services to develop competing products. No benchmarking clause was found. Recorded as a fact with no deduction. It matters before any probe is run
- docs.cartesia.ai's robots.txt carries `Content-Signal: ai-train=yes, search=yes, ai-input=yes`. www.cartesia.ai allows every path. status.cartesia.ai and trust.cartesia.ai serve no robots.txt
- The lead described a REST batch API for Ink 2. The docs say batch `/stt` accepts `ink-whisper` only and mark `ink-2` batch as 'Not available yet'. The listing is named Cartesia Ink and covers both models
- The model page dates `ink-2` 17 September 2026 while the changelog announced Ink 2 in May 2026. The September date is read as the multilingual release, and `lastRelease` uses it
- Cartesia's pricing FAQ quotes $0.39 an hour for Ink on the Scale plan. Plan arithmetic (10,800 credits an hour at $299 for 8 million) gives $0.40
- The privacy policy says the services are designed for users in the United States only, while the docs sell regional deployments in the EU, UK, India and Australia
- The access token page lists an OpenAI-compatible `/audio/transcriptions` endpoint under the STT grant. No reference page for it was found in the docs index
- The `cartesia-mcp` repository has no licence file in its root. The SDK repositories are Apache-2.0
- Both SDK repositories returned 134 stars and 4 open issues from the GitHub API on 9 October. The match was not investigated
- Whether this listing should stand apart from the Cartesia text-to-speech listing is the editor's call. It shares the API host, keys, plans, terms and status page, and has its own models, endpoints, prices and status components

## Weaknesses

- The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls
- Zero data retention is an Enterprise plan setting, and no retention period for other plans was found in the terms, privacy policy or docs
- `ink-2` supports English, French, Hindi, Japanese and Spanish only and has no batch endpoint. `POST /stt` accepts `ink-whisper` only
- No diarisation or speaker labels in the reviewed documentation, and realtime input is raw mono PCM with `encoding` and `sample_rate` required
- The terms of 14 June 2024 forbid automation software (bots) and any robot or scraper that accesses the Services to collect data, which matters before any probe is run
- No SLA was found, and a 429 for exceeding the concurrency limit is documented without a `Retry-After` header or backoff guidance

## What costs an agent a turn today

The notes we give agents before they call it. Each one is a workaround an agent shouldn't need.

- Use `wss://api.cartesia.ai/stt/turns/websocket` with `model=ink-2`, `encoding`, `sample_rate` and `cartesia_version=2026-08-14`. All four are required.
- Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
- Read the final text from `turn.end` only. `transcript` is cumulative within a turn, so joining `turn.update` events duplicates text.
- Send `{"type": "close"}` after the last audio and keep reading until the server closes the socket, or the buffered tail is lost.
- Check `encoding` and `sample_rate` against the source before sending. The docs say the server might not return an error when they are wrong.

## When it's done

Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.

What we couldn't check

  • unchecked: trust.cartesia.ai is drawn by script, so the certifications, the sub-processor list and any disclosure policy or bug bounty were not read
  • unchecked: the duration and scope of the five API-wide incidents on the status history (1 August, 2 and 3 September, 7 and 8 October 2026). Only titles and dates were shown to us
  • unchecked: whether the Free plan asks for a card at sign-up, and how keys are rotated or revoked in the Playground, which is behind a sign-in
  • unchecked: replies on the SDK issue trackers, and weekly download counts on npm and PyPI
  • unchecked: the OpenAPI file at /latest.yml was not fetched whole. The definitions were read as embedded in the reference pages' Markdown twins
  • The Terms of Service forbid using 'automation software (bots)' and 'any robot, spider, crawlers, scraper, or other automatic device' that accesses the Services to monitor or collect data, and forbid using the Services to develop competing products. No benchmarking clause was found. Recorded as a fact with no deduction. It matters before any probe is run
  • docs.cartesia.ai's robots.txt carries Content-Signal: ai-train=yes, search=yes, ai-input=yes. www.cartesia.ai allows every path. status.cartesia.ai and trust.cartesia.ai serve no robots.txt
  • The lead described a REST batch API for Ink 2. The docs say batch /stt accepts ink-whisper only and mark ink-2 batch as 'Not available yet'. The listing is named Cartesia Ink and covers both models
  • The model page dates ink-2 17 September 2026 while the changelog announced Ink 2 in May 2026. The September date is read as the multilingual release, and lastRelease uses it
  • Cartesia's pricing FAQ quotes $0.39 an hour for Ink on the Scale plan. Plan arithmetic (10,800 credits an hour at $299 for 8 million) gives $0.40
  • The privacy policy says the services are designed for users in the United States only, while the docs sell regional deployments in the EU, UK, India and Australia
  • The access token page lists an OpenAI-compatible /audio/transcriptions endpoint under the STT grant. No reference page for it was found in the docs index
  • The cartesia-mcp repository has no licence file in its root. The SDK repositories are Apache-2.0
  • Both SDK repositories returned 134 stars and 4 open issues from the GitHub API on 9 October. The match was not investigated
  • Whether this listing should stand apart from the Cartesia text-to-speech listing is the editor's call. It shares the API host, keys, plans, terms and status page, and has its own models, endpoints, prices and status components

Sources 36

  1. docs index (llms.txt) docs.cartesia.ai · seen 2026-10-09
  2. Ink 2 model page docs.cartesia.ai · seen 2026-10-09
  3. older STT models docs.cartesia.ai · seen 2026-10-09
  4. endpoint comparison docs.cartesia.ai · seen 2026-10-09
  5. realtime (auto) reference, read as the embedded AsyncAPI definition in the Markdown twin docs.cartesia.ai · seen 2026-10-09
  6. realtime (manual) reference, read as the embedded AsyncAPI definition in the Markdown twin docs.cartesia.ai · seen 2026-10-09
  7. batch reference, read as the embedded OpenAPI definition in the Markdown twin docs.cartesia.ai · seen 2026-10-09
  8. turn detection guide docs.cartesia.ai · seen 2026-10-09
  9. audio input guide docs.cartesia.ai · seen 2026-10-09
  10. troubleshooting, realtime auto docs.cartesia.ai · seen 2026-10-09
  11. realtime STT quickstart docs.cartesia.ai · seen 2026-10-09
  12. docs pricing page docs.cartesia.ai · seen 2026-10-09
  13. public pricing page and FAQ cartesia.ai · seen 2026-10-09
  14. API conventions and versioning docs.cartesia.ai · seen 2026-10-09
  15. API errors docs.cartesia.ai · seen 2026-10-09
  16. concurrency and WebSocket limits docs.cartesia.ai · seen 2026-10-09
  17. authentication and access tokens docs.cartesia.ai · seen 2026-10-09
  18. access token reference docs.cartesia.ai · seen 2026-10-09
  19. zero data retention docs.cartesia.ai · seen 2026-10-09
  20. regional endpoints docs.cartesia.ai · seen 2026-10-09
  21. changelog 2026 docs.cartesia.ai · seen 2026-10-09
  22. deprecated models and breaking changes docs.cartesia.ai · seen 2026-10-09
  23. MCP server docs docs.cartesia.ai · seen 2026-10-09
  24. Ink product page cartesia.ai · seen 2026-10-09
  25. Terms of Service, last revised 14 June 2024 cartesia.ai · seen 2026-10-09
  26. Privacy Policy, last revised 14 June 2024 cartesia.ai · seen 2026-10-09
  27. Acceptable Use Policy, last revised 23 July 2025 cartesia.ai · seen 2026-10-09
  28. Data Protection Addendum cartesia.ai · seen 2026-10-09
  29. status page status.cartesia.ai · seen 2026-10-09
  30. status history status.cartesia.ai · seen 2026-10-09
  31. security.txt cartesia.ai · seen 2026-10-09
  32. Python SDK repository, shallow clone, tags and README github.com · seen 2026-10-09
  33. JavaScript SDK repository, shallow clone and tags github.com · seen 2026-10-09
  34. MCP server source, shallow clone github.com · seen 2026-10-09
  35. star and open issue counts api.github.com · seen 2026-10-09
  36. domain registration (RDAP) rdap.org · seen 2026-10-09

Probe metrics

Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.

Pricing & changes

Freemium Freemium Billed in credits. `ink-2` costs 3 credits a second of audio on both realtime endpoints, silence included. `ink-whisper` costs 1 credit a second in realtime and 1 credit per 2 seconds in batch (https://docs.cartesia.ai/pricing). Plans are Free ($0, 20,000 credits a month), Pro ($5, 100,000), Startup ($49, 1.25 million) and Scale ($299, 8 million), so an agent's owner can start without a contract. Commercial use starts at Pro (https://www.cartesia.ai/pricing).

Prices

ItemPriceUnitNote
Ink 2 realtime, Scale plan$0.0067per minute of audio180 credits a minute at $299 for 8M credits. Cartesia quotes $0.39 an hour
Ink 2 realtime, Startup plan$0.0071per minute of audio180 credits a minute at $49 for 1.25M credits
Ink 2 realtime, Pro plan$0.009per minute of audio180 credits a minute at $5 for 100,000 credits
Ink Whisper realtime, Scale plan$0.0022per minute of audio60 credits a minute at $299 for 8M credits
Ink Whisper batch, Scale plan$0.0011per minute of audio30 credits a minute at $299 for 8M credits

Compared across listings on the price index.

Recent changes

  • Latest release

Follow them as a feed at /feeds/tools/cartesia-ink-stt.xml, or this listing's score history at history.json.

Connect

Install

pip install 'cartesia[websockets]'   # or: npm install @cartesia/cartesia-js

Claude Code

claude mcp add --transport http --scope user cartesia https://mcp.cartesia.ai/mcp
Similar toolGrade ScoreShared capabilitiesx402
Amazon Transcribe Amazon Web ServicesBB73.4speech.stt speech.streaming speech.batch speech.languagesno
Azure AI Speech speech-to-text Microsoft AzureBB73speech.stt speech.streaming speech.batch speech.languagesno
OpenAI Speech to Text OpenAIBB72.4speech.stt speech.streaming speech.batch speech.languagesno
Deepgram Speech-to-Text (Nova-3, Flux) DeepgramBB70.3speech.stt speech.streaming speech.batch speech.languagesno
Google Cloud Speech-to-Text Google CloudBB70.2speech.stt speech.streaming speech.batch speech.languagesno
Gladia Speech-to-Text API + MCP GladiaB69.5speech.stt speech.streaming speech.batch speech.languagesno

Machine-readable

Verify this listing

For the vendor

Is this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.

  1. Add the badge or a link

    Cartesia Ink on Anchor Terminal, B, 67.9/100
    On a light page
    On a dark page
    <a href="https://www.anchorterminal.com/tools/cartesia-ink-stt"><img src="https://www.anchorterminal.com/badges/cartesia-ink-stt.svg" alt="Cartesia Ink on Anchor Terminal" height="20"></a>
    [![Cartesia Ink on Anchor Terminal](https://www.anchorterminal.com/badges/cartesia-ink-stt.svg)](https://www.anchorterminal.com/tools/cartesia-ink-stt)

    It counts on a page on cartesia.ai or one of its subdomains, or the README of github.com/cartesia-ai/cartesia-python.

  2. Tell us where it is

    We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.

Agents send the same to POST /api/v1/verify as {"slug": "cartesia-ink-stt", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.