# Cerebrium (slim) > Cerebrium is a serverless platform for running your own models and code on GPUs and CPUs. A CLI packages code into containers served as REST, streaming and WebSocket endpoints, managed through a REST API. - Full: https://www.anchorterminal.com/tools/cerebrium.md (~7,400 tokens) · this version ~1,580 tokens · JSON https://www.anchorterminal.com/tools/cerebrium.json · canonical https://www.anchorterminal.com/tools/cerebrium - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-08 **C · 55.3/100 · rank #512 of 722 · #8 in GPU & serverless compute · not agent-ready · confidence medium** Assessment: Per-second GPU prices are public, a 94-operation OpenAPI spec covers the management API, and service account tokens expire and are limited to named projects. Deployed endpoints are callable without a token unless `disable_auth = false` is set, and no request rate limits, 429 handling or SLA were found in the reviewed documentation. ## Facts - Kind: Model platform · vendor: Cerebrium Inc. · category: GPU & serverless compute · legal entity: Cerebrium Inc. · provenance 80/100 - Endpoint: `https://rest.cerebrium.ai` (HTTP) - Auth: API key · pricing: Freemium · x402: no · licence: Proprietary service under Cerebrium's terms of service. The CLI is MIT - Probe metrics: not measured yet (probes haven't run) - Free tier: Hobby plan, $0 a month plus compute, 3 apps, 3 seats, 5 concurrent GPUs. First 100 GB of storage free. No free compute allowance found on the pricing page - GPUs: T4, L4, A10, L40s, A100 40 GB and 80 GB, RTX PRO 6000, H100, H200, B200, up to 8 per app, plus AWS Inferentia2 and Trainium. `compute` accepts a preference list of up to 5 types - Scale to zero: `min_replicas = 0` by default, `cooldown` in seconds before scale-down, `max_replicas` as the ceiling - Cold start: Cerebrium's home page claims 2 to 4 second cold starts. Cold-start time is not billed, model initialisation is - Timeouts: `response_grace_period` defaults to 15 minutes. Async runs up to 12 hours - Billing basis: Per second for GPU, CPU and memory, including builds. `protected` compute tier at 2x the listed rate - Endpoints: REST, streaming, WebSocket, OpenAI-compatible, async with webhook forwarding, custom ASGI servers and Dockerfiles - Management API: https://rest.cerebrium.ai, OpenAPI 3.0, 94 operations, Bearer service account token - Regions: us-east-1, us-central1, eu-north1 (Finland), eu-north-1 (Stockholm). us-west-2, eu-west-2 and ap-south-1 on request. Multi-region deployment is in beta - Compliance: SOC 2 Type II announced 8 July 2026, HIPAA BAA and DPA on request, per the vendor - Prices: B200 180 GB $6.01 per GPU-hour; H200 141 GB $4.20 per GPU-hour; H100 80 GB $3.40 per GPU-hour; RTX PRO 6000 96 GB $2.50 per GPU-hour; A100 80 GB $2.10 per GPU-hour; A100 40 GB $2 per GPU-hour; L40s 48 GB $1.95 per GPU-hour; A10 24 GB $1.10 per GPU-hour; L4 24 GB $0.80 per GPU-hour; T4 16 GB $0.59 per GPU-hour; CPU-only compute $0.0236 per vCPU-hour; Persistent storage $0.05 per GB per month; Standard plan $100 per month (plan) - Scores: Reliability 48, Performance pending, Schema & documentation 70, Agent ergonomics 49, Security & auth 60, Payments & pricing 30, Task success pending, Maintenance & community 75, Transparency & trust 63 · total over the 7 assessed categories - Why: Reliability, Graded as a hosted service. · Schema & documentation, Public OpenAPI 3.0 spec for the management API with 94 operations on 73 paths. · Agent ergonomics, No field selection. · Security & auth, Service account tokens expire after up to one year (or never, if set to 0), are limited to 1 to 50 named projects and are revoked by deletin… · Payments & pricing, No machine payment protocol (0). · Maintenance & community, CLI 2.9.0 on 16 September 2026, 22 days before the check (30). · Transparency & trust, Closed service with published terms. - Sources: 21, open questions: 6, both in the full twin - Capabilities: compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers - JSON: https://www.anchorterminal.com/api/v1/tools/cerebrium.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/cerebrium.svg` or a link to https://www.anchorterminal.com/tools/cerebrium from a page on cerebrium.ai or one of its subdomains, or the README of github.com/CerebriumAI/cerebrium, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Set `disable_auth = false` in `cerebrium.toml` before deploying. The default leaves the endpoint callable by anyone with the URL 2. Authenticate headless with `CEREBRIUM_SERVICE_ACCOUNT_TOKEN`. `cerebrium login` opens a browser 3. Raise `response_grace_period` for long work. It defaults to 15 minutes and async runs stop at 12 hours 4. Send `?async=true` to get a `run_id` with HTTP 202, and add `webhookEndpoint` because async calls return no result to the caller 5. Check the plan before choosing hardware. A100, H100, H200, B200 and RTX PRO 6000 need Standard, and `protected` compute bills at twice the listed rate ## Connect ```bash pip install cerebrium && cerebrium login ``` ```bash curl --location --request POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/{function}' \ --header 'Authorization: Bearer ' \ --header 'Content-Type: application/json' \ --data '{"function_param": "data"}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Modal | B | 63.6 | compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers | https://www.anchorterminal.com/tools/modal.min.md | | Beam | C | 55.5 | compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers | https://www.anchorterminal.com/tools/beam.min.md | | Runpod | D | 53.5 | compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers | https://www.anchorterminal.com/tools/runpod.min.md | | Baseten | B | 66.5 | compute.gpu, compute.endpoints, compute.serverless, compute.containers | https://www.anchorterminal.com/tools/baseten.min.md | | Replicate Deployments | B | 63.6 | compute.gpu, compute.endpoints, compute.serverless, compute.containers | https://www.anchorterminal.com/tools/replicate-deploy.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)