# Nebius AI Cloud (slim) > Nebius AI Cloud rents NVIDIA GPU virtual machines and InfiniBand clusters, with managed Kubernetes, Slurm and Serverless AI jobs and endpoints for containers. Resources are managed through REST and gRPC APIs, a CLI, a Terraform provider and SDKs. - Full: https://www.anchorterminal.com/tools/nebius-ai-cloud.md (~8,500 tokens) · this version ~1,930 tokens · JSON https://www.anchorterminal.com/tools/nebius-ai-cloud.json · canonical https://www.anchorterminal.com/tools/nebius-ai-cloud - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **B · 67.2/100 · rank #240 of 842 · #1 in GPU & serverless compute · not agent-ready · confidence medium** Assessment: One API definition generates the REST and gRPC interfaces, the CLI, Terraform provider and three SDKs, with a 602-operation OpenAPI document, `X-Idempotency-Key` and role-scoped service accounts. The status page lists 14 major incidents between 14 July and 8 October 2026, no request rate limits were found, and signup needs a browser and a card. ## Facts - Kind: HTTP API · vendor: Nebius · category: GPU & serverless compute · legal entity: Nebius B.V. · provenance 83/100 - Endpoint: `https://api.nebius.cloud` (HTTP, stdio) - Auth: OAuth or key · pricing: Pay per use · x402: no · licence: Proprietary service under the Nebius Services Agreement. The API definitions, the Go, Python and JavaScript SDKs and the MCP server on GitHub are MIT - Probe metrics: not measured yet (probes haven't run) - Free tier: None found. A card added at signup is charged $25, which goes to the account balance. Promo codes are issued in promotions - GPUs: NVIDIA B300, B200, H200 and H100 (NVLink), RTX PRO 6000 and L40S, plus CPU-only AMD and Intel platforms. GB300 and GB200 NVL72 racks through sales - Interfaces: REST at `https://api.nebius.cloud` (OpenAPI 3.0.3, 602 operations), gRPC on per-service hosts such as `compute.api.nebius.cloud:443`, the `nebius` CLI, a Terraform provider and SDKs for Go, Python and JavaScript, all generated from one API - Serverless AI: Devlabs, jobs and endpoints that run a container image on a managed container VM. Endpoints get a managed HTTPS URL and optional token authentication (https://docs.nebius.com/serverless/overview) - Scale to zero: None automatic. An endpoint is stopped and started by a call, and a stopped endpoint bills for neither compute nor storage. No autoscaling found - Billing basis: Hourly prices charged by the second. Volumes bill per GiB for 730 hours whether attached or not - Idempotency and retries: `X-Idempotency-Key` on modifying calls, `resourceVersion` for concurrent updates, and a `retry_type` on error details (https://github.com/nebius/api) - Rate limits: No request rate limits found. Default quotas per region include 12 regular and 8 preemptible GPU VMs and 32 H200 GPUs in `eu-north1` (https://docs.nebius.com/compute/resources/quotas-limits) - SLA: 99.5 per cent a month per virtual machine, with credits of 10, 15 or 30 per cent. Serverless AI is not on the list of services with a service level (https://docs.nebius.com/legal/sla-levels) - Regions: Nine public regions in Finland, France (two), Spain, Israel, the United Kingdom (two) and the United States (Missouri and Minnesota), plus a private region in Iceland - MCP servers: A keyless docs search server at `https://docs.nebius.com/mcp`, and a beta local server (`nebius/mcp-server`, four tools) that runs CLI commands with a safe mode on by default - Certifications: SOC 2 Type II with HIPAA, SOC 3, ISO 27001, 27018 and 22301, and CSA STAR Level 1 per https://nebius.com/trust-center - Audit: Audit Logs (preview, free) with control-plane and data-plane events, an API at `/audit/v2/audit-events` and export to Object Storage - Prices: NVIDIA H200 NVLink, on demand $5.40 per GPU-hour; NVIDIA H100 NVLink, on demand $4.50 per GPU-hour; NVIDIA B200 NVLink, on demand $8.50 per GPU-hour; NVIDIA B300 NVLink, on demand $9.50 per GPU-hour; NVIDIA RTX PRO 6000, on demand $1.80 per GPU-hour; NVIDIA L40S, GPU only $1.35 per GPU-hour; NVIDIA H200 NVLink, preemptible $0.79 per GPU-hour; Network SSD disk $0.071 per GB per month - Scores: Reliability 57, Performance pending, Schema & documentation 78, Agent ergonomics 77, Security & auth 81, Payments & pricing 20, Task success pending, Maintenance & community 82, Transparency & trust 77 · total over the 7 assessed categories - Why: Reliability, Graded as a hosted service on the REST API. · Schema & documentation, OpenAPI 3.0.3 at api.nebius.cloud/openapi.json with 602 operations on 443 paths and 1,051 schemas, plus the protobuf definitions in… · Agent ergonomics, Lists take `pageSize` and `pageToken`, gets take a `view` and there are get-by-name calls; no field selection (15 of 25). · Security & auth, User tokens last 12 hours. · Payments & pricing, No machine payment protocol (0). · Maintenance & community, Python SDK 0.6.20 on 7 October 2026 and CLI 0.12.287 on 6 October (30). · Transparency & trust, Closed service under a Services Agreement published 15 September 2026 and effective 28 September, with nine earlier versions archived; the A… - Sources: 31, open questions: 7, both in the full twin - Capabilities: compute.gpu, compute.endpoints, compute.batch, compute.containers - JSON: https://www.anchorterminal.com/api/v1/tools/nebius-ai-cloud.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/nebius-ai-cloud.svg` or a link to https://www.anchorterminal.com/tools/nebius-ai-cloud from a page on nebius.com or one of its subdomains, or the README of github.com/nebius/api, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Use a service account with an authorised key, then exchange a five-minute RS256 JWT at `https://auth.eu.nebius.com/oauth2/token/exchange` for a 12-hour Bearer token 2. Send `X-Idempotency-Key` with a random UUID on every create, update and delete, since a 504 can follow a call that succeeded 3. Poll the returned operation (`/ai/v1/endpoints/operations/{id}`) until `status` is set; concurrent operations on one resource are not supported 4. Stop or delete endpoints when idle. A stopped endpoint bills nothing, a stopped Devlab or VM still bills for its disk 5. Check region support first. Serverless AI is absent from `eu-south1` and `us-north1`, and each GPU platform exists in one to four regions 6. Run the beta `nebius/mcp-server` with safe mode on (the default); `nebius_cli_execute` can run any CLI command when `SAFE_MODE=false` ## Connect ```bash curl -sSL https://artifacts.nebius.cloud/cli/install.sh | bash ``` ```bash curl --request GET --url 'https://api.nebius.cloud/iam/v1/profiles' --header 'Authorization: Bearer ' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/nebius-ai-cloud ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | Modal | B | 63.6 | compute.gpu, compute.endpoints, compute.batch, compute.containers | https://www.anchorterminal.com/tools/modal.min.md | | Verda | B | 62.3 | compute.gpu, compute.endpoints, compute.batch, compute.containers | https://www.anchorterminal.com/tools/verda.min.md | | CoreWeave | C | 61.5 | compute.gpu, compute.containers, compute.endpoints, compute.batch | https://www.anchorterminal.com/tools/coreweave.min.md | | Northflank | C | 61.3 | compute.gpu, compute.containers, compute.batch, compute.endpoints | https://www.anchorterminal.com/tools/northflank.min.md | | Beam | C | 55.5 | compute.gpu, compute.endpoints, compute.batch, compute.containers | https://www.anchorterminal.com/tools/beam.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)