# Gemini Live API (slim) > Google's Live API runs real-time spoken conversations with Gemini audio-to-audio models over a stateful WebSocket. It takes audio, images and text, speaks back, and supports interruptions, function calling and Google Search grounding. - Full: https://www.anchorterminal.com/tools/gemini-live.md (~9,100 tokens) · this version ~2,030 tokens · JSON https://www.anchorterminal.com/tools/gemini-live.json · canonical https://www.anchorterminal.com/tools/gemini-live - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 57.1/100 · rank #560 of 842 · #7 in Conversational voice agents · not agent-ready · confidence medium** Assessment: A speech-to-speech API with per-token prices published, a free tier and a generally available model, `gemini-3.8-live`, since 15 September 2026. The documented raw WebSocket connection carries the API key in the URL, connections reset about every 10 minutes, and no SLA or readable incident history was found for the Developer API. ## Facts - Kind: HTTP API · vendor: Google · category: Conversational voice agents · legal entity: Google LLC · provenance 89/100 - Endpoint: `https://generativelanguage.googleapis.com/v1beta` (HTTP) - Auth: API key · pricing: Freemium · x402: no · licence: Proprietary service under the Gemini API Additional Terms of Service. The Python and JavaScript SDKs are Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - Architecture: Speech-to-speech. Native audio models take audio in and speak audio out over one stateful WebSocket, with optional text transcripts of both sides - Endpoint: `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent`, and `BidiGenerateContentConstrained` for ephemeral tokens - Models: `gemini-3.8-live` (stable, default), `gemini-3.8-live-extended-thinking`, `gemini-3.1-flash-live-preview` and `gemini-2.5-flash-native-audio-preview-12-2025` (both with an earliest shutdown of 17 November 2026). `gemini-3.5-live-translate-preview` and `gemini-3.5-transcribe-live` use the same API - Audio: Input raw 16-bit PCM at 16 kHz, images as JPEG at up to 1 frame a second, and text. Output raw 16-bit PCM at 24 kHz. Audio is the only response modality on native audio models - Languages: The overview page says 70 and the capabilities guide lists 99. The model detects the spoken language without a language code - Latency: No figure published in the docs read. Google describes the API as low latency - Interruptions: Automatic voice activity detection by default, with start and end sensitivity, prefix padding and silence duration settings. Hybrid and manual modes. The server sends `interrupted: true` when the user barges in - Tool calling: Function declarations and Google Search grounding. Non-blocking calls by default on `gemini-3.8-live`, with `SILENT`, `WHEN_IDLE` and `INTERRUPT` scheduling. No code execution, URL context or Google Maps. The client must send tool responses itself - Telephony: None built in. Partner integrations such as Voximplant, LiveKit, Pipecat and Agora connect calls - Session limits: Audio sessions 15 minutes and audio with video 2 minutes without context compression. A connection lasts about 10 minutes. Resumption handles are valid 2 hours. Context window 128k tokens on native audio models - Credentials: API key from AI Studio, an authorisation key bound to a service account by default since 28 May 2026. Ephemeral tokens (preview) from `POST /v1beta/auth_tokens` with `uses`, `expireTime`, `newSessionExpireTime` and `liveConnectConstraints` - Free tier: Live models are free of charge on the free tier, where content is used to improve Google products. Limits are shown in AI Studio - Rate limits: Per project and usage tier, shown in AI Studio. Public figures are the spend limits of $10, $50 and $200 per 10 minutes on Tiers 1 to 3. No concurrent session figure in the public docs - Data retention: 55 days for abuse monitoring. Up to 24 hours of session state when resumption is on. 30 days for Google Search grounding. Paid-tier content isn't used to improve Google products - SDKs: `google-genai` v2.29.0 for Python and `@google/genai` v2.28.0 for JavaScript, both tagged 7 October 2026, Apache-2.0 - Prices: gemini-3.8-live audio input $0.005 per minute of audio; gemini-3.8-live audio output $0.018 per minute of audio; gemini-3.8-live text input $0.75 per 1M tokens; gemini-3.8-live text output $4.50 per 1M tokens; Google Search grounding $14 per 1,000 requests - Scores: Reliability 41, Performance pending, Schema & documentation 66, Agent ergonomics 69, Security & auth 48, Payments & pricing 40, Task success pending, Maintenance & community 83, Transparency & trust 72 · total over the 7 assessed categories - Why: Reliability, Graded as a hosted API on the Gemini Developer API. · Schema & documentation, No AsyncAPI or similar file for the WebSocket messages was found. · Agent ergonomics, Scored as an API. · Security & auth, Keys created since 28 May 2026 are authorisation keys bound to a service account and restricted to the Gemini API, with IP and origin restri… · Payments & pricing, No x402, MPP or L402 (0 of 40). · Maintenance & community, `gemini-3.8-live` went generally available on 15 September 2026, 23 days before the check (30). · Transparency & trust, Closed service with clear terms, SDKs under Apache-2.0 (15 of 30). - Sources: 29, open questions: 9, both in the full twin - Capabilities: voice.agent, voice.speech-to-speech, voice.tools - JSON: https://www.anchorterminal.com/api/v1/tools/gemini-live.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/gemini-live.svg` or a link to https://www.anchorterminal.com/tools/gemini-live from a page on google.com or ai.google.dev or one of their subdomains, or the README of github.com/google-gemini/gemini-live-api-examples, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Enable `sessionResumption` and keep the newest handle. Connections end after about 10 minutes, and handles stay valid for 2 hours. 2. Set `contextWindowCompression` with a sliding window. Without it audio sessions stop at 15 minutes and audio with video at 2 minutes, and every turn re-bills the whole context. 3. On `gemini-3.8-live` function calls are non-blocking by default. Set `behavior: BLOCKING` if the model must wait for the tool response. 4. Send 16 kHz 16-bit PCM in 20 to 40 ms chunks and discard buffered playback when `interrupted` is true. 5. Keep the API key on a server and send it through the SDK. Give browsers an ephemeral token from `POST /v1beta/auth_tokens`, locked with `liveConnectConstraints`. ## Connect ```bash pip install google-genai # or: npm i @google/genai ``` ```bash curl -X POST "https://generativelanguage.googleapis.com/v1beta/auth_tokens" \ -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \ -d '{"uses": 1, "liveConnectConstraints": {"model": "models/gemini-3.8-live", "config": {"sessionResumption": {}, "responseModalities": ["AUDIO"]}}}' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/gemini-live ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Retell AI API + MCP | B | 69.1 | voice.agent, voice.speech-to-speech, voice.tools | https://www.anchorterminal.com/tools/retell-ai.min.md | | Vapi API + MCP | B | 63.5 | voice.agent, voice.speech-to-speech, voice.tools | https://www.anchorterminal.com/tools/vapi.min.md | | Ultravox Realtime API | C | 58.5 | voice.agent, voice.speech-to-speech, voice.tools | https://www.anchorterminal.com/tools/ultravox.min.md | | Hume EVI (Empathic Voice Interface) | C | 56.7 | voice.agent, voice.speech-to-speech, voice.tools | https://www.anchorterminal.com/tools/hume-evi.min.md | | Bolna API + MCP | D | 52.6 | voice.agent, voice.speech-to-speech, voice.tools | https://www.anchorterminal.com/tools/bolna.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)