Head to head · Speech-to-text · October 2026 research run
Cartesia Ink vs Groq Speech-to-Text
Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 5 of 7 scored categories. Cartesia Ink leads on schema & documentation and maintenance & community. Both do speech-to-text.
Best speech-to-text APIs for AI agents · All 91 stt comparisons
Which one, for what
Good for Suited to voice agents that need turn detection and transcription from one model in English, French, Hindi, Japanese or Spanish, and to teams already using Cartesia text-to-speech.
Ahead on
- Schema & documentation, 89 against 58
- Maintenance & community, 85 against 68
Watch for
The terms let Cartesia train models on inputs and outputs unless otherwise agreed. Opting out is a form in the Playground's data controls
Good for Suited to cheap, fast transcription of recorded files in many languages, and to agents that already hold a Groq key or use an OpenAI-compatible client.
Ahead on
- Reliability, 90 against 73
- Security & auth, 79 against 52
- Payments & pricing, 40 against 35
- Transparency & trust, 85 against 67
Also in its favour
- Agent-ready, a grade of BB or better
- Free to start without a card
Watch for
No streaming or realtime endpoint and no diarisation in the reviewed documentation
Score by category
| Category | Weight this run | Cartesia Ink | Groq Speech-to-Text | Edge |
|---|---|---|---|---|
| Reliability | 16%20 | 73 | 90 | Groq Speech-to-Text +17 |
| Performance | 10%pending | pending | pending | not scored in this run |
| Schema & documentation | 13%16.2 | 89 | 58 | Cartesia Ink +31 |
| Agent ergonomics | 13%16.2 | 74 | 75 | Groq Speech-to-Text +1 |
| Security & auth | 14%17.5 | 52 | 79 | Groq Speech-to-Text +27 |
| Payments & pricing | 10%12.5 | 35 | 40 | Groq Speech-to-Text +5 |
| Task success | 10%pending | pending | pending | not scored in this run |
| Maintenance & community | 7%8.8 | 85 | 68 | Cartesia Ink +17 |
| Transparency & trust | 7%8.8 | 67 | 85 | Groq Speech-to-Text +18 |
| Negative events | ≤15 | 0 | 0 | |
| Total | 67.9 · B | 71.8 · BB |
Facts side by side
| Fact | Cartesia Ink | Groq Speech-to-Text |
|---|---|---|
| Kind | Model API | Model API |
| Vendor | Cartesia | Groq |
| Hosted endpoint | https://api.cartesia.ai | https://api.groq.com/openai/v1 |
| Transports | HTTP, websocket, Streamable HTTP | HTTP |
| Auth | API key | API key |
| Pricing | Freemium | Freemium |
| x402 | no | no |
| Licence | Proprietary hosted service under the Cartesia Terms of Service. The Python and JavaScript SDKs are Apache-2.0 | Proprietary hosted service under the Groq Services Agreement. The SDKs are Apache-2.0 and the Whisper weights are published by OpenAI on Hugging Face |
| Read-only variant documented | no | no |
| llms.txt | yes | yes |
| Last release | 2026-09-17 | 2026-08-26 |
| Terms last updated | 2024-06-14 | 2026-06-22 |
| Privacy policy last updated | 2024-06-14 | 2025-11-12 |
| Customer content may train models | yes, with an opt-out | not found in the text |
| Terms restrict automated access | yes | not found in the text |
| Terms restrict benchmarking | yes | yes |
| Terms or service can change without notice | not found in the text | not found in the text |
| Arbitration or class-action waiver | yes | not found in the text |
| Popularity | 134 stars | 621 stars |
Verdicts
Cartesia Ink
Ink 2 streams transcripts with turn detection built into the model, documented by an AsyncAPI file with typed ranges and structured errors. The terms, last revised 14 June 2024, let Cartesia train on inputs and outputs unless the customer opts out, and zero data retention is an Enterprise setting. Ink 2 has no batch endpoint and no diarisation.
Groq Speech-to-Text
Whisper Large v3 Turbo costs $0.04 an audio hour and Whisper Large v3 $0.111, with a no-card free plan and zero data retention as a self-serve setting. There is no streaming endpoint, no diarisation and no subtitle output, and uploads stop at 25 MB on the free plan and 100 MB on the Developer plan.
Before you call either
Cartesia Ink
- Use
wss://api.cartesia.ai/stt/turns/websocketwithmodel=ink-2,encoding,sample_rateandcartesia_version=2026-08-14. All four are required. - Send raw mono audio in chunks of about 100 ms at the speed it was spoken. Pushing a whole file into the socket can return an internal server error.
- Read the final text from
turn.endonly.transcriptis cumulative within a turn, so joiningturn.updateevents duplicates text. - Send
{"type": "close"}after the last audio and keep reading until the server closes the socket, or the buffered tail is lost. - Check
encodingandsample_rateagainst the source before sending. The docs say the server might not return an error when they are wrong.
Groq Speech-to-Text
- Send
whisper-large-v3-turbofor transcription andwhisper-large-v3for translation to English. The translations endpoint does not accept Turbo. - Pass
urlinstead offilefor audio over 25 MB, and split anything over the plan's size limit into overlapping chunks before sending. - Set
response_formattoverbose_jsonbefore asking fortimestamp_granularities[]. Word timestamps add latency, segment timestamps do not. - Every request is billed as at least 10 seconds of audio, so join very short clips where the task allows.
- Read
retry-afteron a 429 and back off. Audio limits count seconds an hour and a day as well as requests.
Questions
Which is better for AI agents, Cartesia Ink or Groq Speech-to-Text?
Groq Speech-to-Text scores 71.8 (BB) on agent readiness against Cartesia Ink's 67.9 (B), and leads in 5 of 7 scored categories. Cartesia Ink leads on schema & documentation and maintenance & community.
Do Cartesia Ink and Groq Speech-to-Text need an API key?
Both need an API key.
Can an agent call Cartesia Ink and Groq Speech-to-Text without installing anything?
Yes. Cartesia Ink has a hosted endpoint at https://api.cartesia.ai and Groq Speech-to-Text at https://api.groq.com/openai/v1.
Other comparisons with Cartesia Ink or Groq Speech-to-Text
- Amazon Transcribe vs Cartesia Ink
- Amazon Transcribe vs Groq Speech-to-Text
- AssemblyAI Speech-to-Text (Universal) vs Cartesia Ink
- AssemblyAI Speech-to-Text (Universal) vs Groq Speech-to-Text
- Azure AI Speech speech-to-text vs Cartesia Ink
- Azure AI Speech speech-to-text vs Groq Speech-to-Text
- Cartesia Ink vs Deepgram Speech-to-Text (Nova-3, Flux)
- Cartesia Ink vs ElevenLabs Scribe Speech to Text API
- Cartesia Ink vs Gladia Speech-to-Text API + MCP
- Cartesia Ink vs Google Cloud Speech-to-Text
- Cartesia Ink vs Mistral Voxtral Transcribe
- Cartesia Ink vs OpenAI Speech to Text
- Cartesia Ink vs Rev AI Speech-to-Text API
- Cartesia Ink vs Soniox Speech-to-Text
- Cartesia Ink vs Speechmatics Speech-to-Text
- Deepgram Speech-to-Text (Nova-3, Flux) vs Groq Speech-to-Text
- ElevenLabs Scribe Speech to Text API vs Groq Speech-to-Text
- Gladia Speech-to-Text API + MCP vs Groq Speech-to-Text
- Google Cloud Speech-to-Text vs Groq Speech-to-Text
- Groq Speech-to-Text vs Mistral Voxtral Transcribe
- Groq Speech-to-Text vs OpenAI Speech to Text
- Groq Speech-to-Text vs Rev AI Speech-to-Text API
- Groq Speech-to-Text vs Soniox Speech-to-Text
- Groq Speech-to-Text vs Speechmatics Speech-to-Text
Machine-readable
- This page as Markdown
/compare/cartesia-ink-stt-vs-groq-speech-to-text.md· slim.min.md· JSON.json(or sendAccept: text/markdown) - Each listing in full
/api/v1/tools/cartesia-ink-stt.json·/api/v1/tools/groq-speech-to-text.json - From a terminal
anchor compare cartesia-ink-stt groq-speech-to-text(the CLI) - Over MCP
compare_tools {"a": "cartesia-ink-stt", "b": "groq-speech-to-text"}at/mcp, no key