# OpenAI Speech to Text (slim) > OpenAI's speech-to-text API. It transcribes uploaded audio files through /v1/audio/transcriptions, translates recordings into English through /v1/audio/translations, and transcribes live audio in Realtime transcription sessions over WebSocket or WebRTC. - Full: https://www.anchorterminal.com/tools/openai-speech-to-text.md (~7,800 tokens) · this version ~1,730 tokens · JSON https://www.anchorterminal.com/tools/openai-speech-to-text.json · canonical https://www.anchorterminal.com/tools/openai-speech-to-text - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-10 **BB · 72.4/100 · rank #106 of 950 · #3 in Speech-to-text · agent-ready · confidence medium** Assessment: `gpt-transcribe` costs $0.0045 an audio minute, and the audio endpoints keep no abuse-monitoring logs or application state. Speaker labels, timestamps, subtitles and translation exist only on `whisper-1` and `gpt-4o-transcribe-diarize`, which shut down on 26 February 2027 with no named replacement for those functions. ## Facts - Kind: Model API · vendor: OpenAI · category: Speech-to-text · legal entity: not named · provenance 59/100 - Endpoint: `https://api.openai.com/v1` (HTTP, websocket) - Auth: API key · pricing: Pay per use · x402: no · licence: Proprietary hosted service. The service terms were not read (see open questions). The Python SDK is Apache-2.0 and the OpenAPI document is MIT - Probe metrics: not measured yet (probes haven't run) - Models: `gpt-transcribe` (files and committed Realtime turns) and `gpt-live-transcribe` (live audio). `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe` and `gpt-4o-transcribe-diarize` are deprecated and shut down on 26 February 2027 - Languages: `languages` takes ISO 639-1 codes, selected ISO 639-3 codes and regional `zh` codes. No count is given for `gpt-transcribe`. The guide says Whisper supports 98 languages - Max audio: 25 MB a file. Longer recordings are split by the caller - Input: Multipart `file` in mp3, mp4, mpeg, mpga, m4a, wav or webm per the guide. The OpenAPI document also lists flac and ogg - Output: `json` with text and detected languages, or `text`. `verbose_json`, `srt` and `vtt` on `whisper-1` only, `diarized_json` on `gpt-4o-transcribe-diarize` only - Streaming: `stream=true` on file transcription (not `whisper-1`), and Realtime transcription sessions over WebSocket or WebRTC for live audio - Diarisation: `gpt-4o-transcribe-diarize` only, with up to four known speaker references. Not available in Realtime sessions. The model is deprecated - Timestamps: Word and segment timestamps on `whisper-1` only - Translation: Into English only, through `/v1/audio/translations` with `whisper-1` - Rate limits: `gpt-transcribe` 5,000 requests a minute on Build, 10,000 on Launch and 30,000 on Grow - Batch: The `gpt-transcribe` model page marks the Batch API as not supported - Trains on API data: No, unless the customer opts in, per the data controls page - Data retention: No abuse-monitoring retention and no application state for `/v1/audio/transcriptions` and `/v1/audio/translations` - Data location: Regional storage in ten regions through project settings and prefixed hosts such as `eu.api.openai.com`, on approval through sales. Regional processing in the United States and Europe - Prices: gpt-transcribe $0.0045 per minute of audio; gpt-live-transcribe (live audio) $0.017 per minute of audio; whisper-1 (deprecated) $0.006 per minute of audio; gpt-4o-transcribe-diarize (deprecated, estimated from token prices) $0.006 per minute of audio - Scores: Reliability 80, Performance pending, Schema & documentation 88, Agent ergonomics 78, Security & auth 86, Payments & pricing 20, Task success pending, Maintenance & community 75, Transparency & trust 61 · total over the 7 assessed categories - Why: Reliability, Hosted reading. · Schema & documentation, Model reading. · Agent ergonomics, API reading, as for the other speech-to-text listings. · Security & auth, Model reading. · Payments & pricing, No x402, MPP or L402 (0). · Maintenance & community, Model reading. · Transparency & trust, The hosted models are closed. - Sources: 18, open questions: 15, both in the full twin - Capabilities: speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages - JSON: https://www.anchorterminal.com/api/v1/tools/openai-speech-to-text.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/openai-speech-to-text.svg` or a link to https://www.anchorterminal.com/tools/openai-speech-to-text from a page on openai.com or one of its subdomains, or the README of github.com/openai/openai-python, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send `gpt-transcribe` to `POST /v1/audio/transcriptions` for recorded files. Use `languages` (a list), not `language`, and never send both. 2. Keep each upload at 25 MB or less. Split longer audio between sentences and pass the previous chunk's text in `prompt`. 3. For speaker labels send `gpt-4o-transcribe-diarize` with `response_format=diarized_json` and `chunking_strategy=auto` for audio over 30 seconds. Plan for its shutdown on 26 February 2027. 4. Word timestamps, `srt`, `vtt` and `/v1/audio/translations` need `whisper-1`, which cannot stream and shuts down on the same date. 5. On 429 or 503 wait at least `Retry-After` when present, then back off with jitter. Do not retry `credit_balance_exhausted` or spend-limit errors. ## Connect ```bash pip install openai ``` ```bash curl --request POST \ --url https://api.openai.com/v1/audio/transcriptions \ --header "Authorization: Bearer $OPENAI_API_KEY" \ --header 'Content-Type: multipart/form-data' \ --form file=@/path/to/file/audio.mp3 \ --form model=gpt-transcribe ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Amazon Transcribe | BB | 73.4 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/amazon-transcribe.min.md | | Azure AI Speech speech-to-text | BB | 73 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/azure-speech-to-text.min.md | | Deepgram Speech-to-Text (Nova-3, Flux) | BB | 70.3 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/deepgram-stt.min.md | | Google Cloud Speech-to-Text | BB | 70.2 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/google-speech-to-text.min.md | | Gladia Speech-to-Text API + MCP | B | 69.5 | speech.stt, speech.streaming, speech.batch, speech.diarisation, speech.languages | https://www.anchorterminal.com/tools/gladia-stt.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)