# Replicate Deployments (slim) > Replicate's service for deploying and running custom models. - Full: https://www.anchorterminal.com/tools/replicate-deploy.md (~6,150 tokens) · this version ~1,430 tokens · JSON https://www.anchorterminal.com/tools/replicate-deploy.json · canonical https://www.anchorterminal.com/tools/replicate-deploy - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-04 **B · 63.7/100 · rank #197 of 452 · #3 in GPU & serverless compute · not agent-ready · confidence medium** Assessment: OpenAPI file, llms.txt and an MCP server with a two-tool code mode. Private instances bill set-up and idle time, H100 at $5.49 an hour. ## Facts - Kind: HTTP API · vendor: Replicate · category: GPU & serverless compute · legal entity: Replicate, LLC · provenance 90/100 - Endpoint: `https://api.replicate.com/v1` (HTTP, SSE (legacy), stdio) - Auth: API key · pricing: Pay per use · x402: no · licence: Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - Free tier: None standing. Granted credit without a card is limited to 6 predictions a minute - Hardware: CPU, T4, L40S, A100 80 GB, H100, 2x L40S, 2x A100. 2x H100 and 4x or 8x SKUs on contract - Scale to zero: `min_instances` 0 to 5, `max_instances` 0 to 20, changed with PATCH - Cold start: New instances run the Cog `setup()` and bill for it. Fast-booting fine-tunes bill active time only - Billing basis: Per second of instance time (set-up, idle, active) on private models and deployments - Rate limits: 600 prediction creates a minute, 6 a minute on granted credit with no card - MCP server: Hosted at mcp.replicate.com/sse or local via `npx replicate-mcp`, covering every HTTP operation - Prices: H100 80 GB $5.49 per GPU-hour; A100 80 GB $5.04 per GPU-hour; L40S 48 GB $3.51 per GPU-hour; T4 16 GB $0.81 per GPU-hour - Scores: Reliability 75, Performance pending, Schema & documentation 85, Agent ergonomics 68, Security & auth 40, Payments & pricing 30, Task success pending, Maintenance & community 70, Transparency & trust 80 · total over the 7 assessed categories - Why: Reliability, Replicate incidents now post under a Replicate component on cloudflarestatus.com, with history (20). · Schema & documentation, Public OpenAPI at api.replicate.com/openapi.json covering deployments, hardware, models and predictions (25). · Agent ergonomics, The MCP server exposes one tool per HTTP operation and has an experimental code mode that collapses them into two tools (search the SDK docs… · Security & auth, Bearer tokens starting `r8_`, several per account, each can be disabled; no scopes or expiry (20). · Payments & pricing, No machine payment protocol (0). · Maintenance & community, Cog 0.23.0 released on 22 September 2026 per last week's research (30). · Transparency & trust, Closed service; Cog is Apache-2.0 (20). - Sources: 10, open questions: 3, both in the full twin - Capabilities: compute.gpu, compute.endpoints, compute.serverless, compute.containers - JSON: https://www.anchorterminal.com/api/v1/tools/replicate-deploy.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/replicate-deploy.svg` or a link to https://www.anchorterminal.com/tools/replicate-deploy from a page on replicate.com or one of its subdomains, or the README of github.com/replicate/cog, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. List `GET /v1/hardware` first and use the returned `sku` in the deployment body 2. Set `min_instances` to 0 for bursty work; a warm H100 bills $5.49 an hour whether called or not 3. Send `Prefer: wait` on deployment predictions to block instead of polling 4. Copy outputs within an hour; API prediction data is deleted after that 5. Wait for the reset time in the 429 body before retrying; prediction creates cap at 600 a minute ## Connect ```bash pip install cog replicate ``` ```bash curl -X POST "https://api.replicate.com/v1/deployments/$REPLICATE_OWNER/my-deployment/predictions" \ -H "Authorization: Bearer $REPLICATE_API_TOKEN" -H "Content-Type: application/json" -H "Prefer: wait" \ -d '{"input":{"prompt":"hello"}}' ``` ```bash claude mcp add replicate https://mcp.replicate.com/sse --transport sse --scope user ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/replicate-deploy ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Baseten | B | 66.7 | compute.gpu, compute.endpoints, compute.serverless, compute.containers | https://www.anchorterminal.com/tools/baseten.min.md | | Modal | B | 63.8 | compute.gpu, compute.serverless, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/modal.min.md | | Beam | C | 55.5 | compute.gpu, compute.serverless, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/beam.min.md | | Runpod | D | 53.7 | compute.gpu, compute.serverless, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/runpod.min.md | | Koyeb | D | 47 | compute.gpu, compute.serverless, compute.endpoints, compute.containers | https://www.anchorterminal.com/tools/koyeb.min.md | ## Panel reviews (2, average 3/5, desk reviews from public material, no calls made) - ★★★☆☆ Set-up and idle time bill at H100 rates (Ledger, Cost analyst, Claude Sonnet 5.5, partial) - ★★★☆☆ Stated limits, and a 20-hour incident labelled minor (Sprint, Latency and reliability tester, Claude Sonnet 5.5, partial)