Nebius AI Cloud
by Nebius HTTP API in GPU & serverless compute
Hosted Local
Nebius B.V. · nebius.com since 2004 · status page · who's behind it
Nebius AI Cloud rents NVIDIA GPU virtual machines and InfiniBand clusters, with managed Kubernetes, Slurm and Serverless AI jobs and endpoints for containers. Resources are managed through REST and gRPC APIs, a CLI, a Terraform provider and SDKs.
Good for Teams that want whole GPU VMs or InfiniBand clusters in Europe, the UK, Israel or the US with IAM, Terraform and an SLA, and are content to manage endpoint lifecycles themselves.
Is this your product? Claim this listing or verify it
More from Nebius Nebius Token Factory fine-tuning (Fine-tuning)
Assessment. One API definition generates the REST and gRPC interfaces, the CLI, Terraform provider and three SDKs, with a 602-operation OpenAPI document, X-Idempotency-Key and role-scoped service accounts. The status page lists 14 major incidents between 14 July and 8 October 2026, no request rate limits were found, and signup needs a browser and a card.
Facts
- Transport
- HTTP, stdio
- Endpoint
https://api.nebius.cloud- Auth
- OAuth or key
- Pricing
- Pay per use · Pay per use
- x402
- No
- Licence
- Proprietary service under the Nebius Services Agreement. The API definitions, the Go, Python and JavaScript SDKs and the MCP server on GitHub are MIT
- Packages
pypinebiusnpm@nebius/js-sdkgogithub.com/nebius/gosdk- Source
- github.com/nebius/api
- Docs
- docs.nebius.com/
- llms.txt
- published
- Last release
- npm / week
- 4.1k
- PyPI / week
- 469k
- Free tier
- None found. A card added at signup is charged $25, which goes to the account balance. Promo codes are issued in promotions
- GPUs
- NVIDIA B300, B200, H200 and H100 (NVLink), RTX PRO 6000 and L40S, plus CPU-only AMD and Intel platforms. GB300 and GB200 NVL72 racks through sales
- Interfaces
- REST at
https://api.nebius.cloud(OpenAPI 3.0.3, 602 operations), gRPC on per-service hosts such ascompute.api.nebius.cloud:443, thenebiusCLI, a Terraform provider and SDKs for Go, Python and JavaScript, all generated from one API - Serverless AI
- Devlabs, jobs and endpoints that run a container image on a managed container VM. Endpoints get a managed HTTPS URL and optional token authentication (https://docs.nebius.com/serverless/overview)
- Scale to zero
- None automatic. An endpoint is stopped and started by a call, and a stopped endpoint bills for neither compute nor storage. No autoscaling found
- Billing basis
- Hourly prices charged by the second. Volumes bill per GiB for 730 hours whether attached or not
- Idempotency and retries
X-Idempotency-Keyon modifying calls,resourceVersionfor concurrent updates, and aretry_typeon error details (https://github.com/nebius/api)- Rate limits
- No request rate limits found. Default quotas per region include 12 regular and 8 preemptible GPU VMs and 32 H200 GPUs in
eu-north1(https://docs.nebius.com/compute/resources/quotas-limits) - SLA
- 99.5 per cent a month per virtual machine, with credits of 10, 15 or 30 per cent. Serverless AI is not on the list of services with a service level (https://docs.nebius.com/legal/sla-levels)
- Regions
- Nine public regions in Finland, France (two), Spain, Israel, the United Kingdom (two) and the United States (Missouri and Minnesota), plus a private region in Iceland
- MCP servers
- A keyless docs search server at
https://docs.nebius.com/mcp, and a beta local server (nebius/mcp-server, four tools) that runs CLI commands with a safe mode on by default - Certifications
- SOC 2 Type II with HIPAA, SOC 3, ISO 27001, 27018 and 22301, and CSA STAR Level 1 per https://nebius.com/trust-center
- Audit
- Audit Logs (preview, free) with control-plane and data-plane events, an API at
/audit/v2/audit-eventsand export to Object Storage - Capabilities
- compute.gpu compute.endpoints compute.batch compute.containers
Facts verified 2026-10-08 from vendor docs, repositories and package registries. JSON · Markdown
Strengths
- OpenAPI 3.0.3 document at
https://api.nebius.cloud/openapi.jsonwith 602 operations, generated from the same protobuf definitions as the gRPC API, CLI, Terraform provider and SDKs X-Idempotency-Keyheader for modifying calls, and aretry_typefield on errors that says whether to retry the call- Service accounts sign in with an uploaded RSA key and receive 12-hour tokens, with roles granted per tenant, project or resource
- Per-second billing with public prices, and a 99.5 per cent monthly uptime commitment per virtual machine
- llms.txt, every docs page as Markdown, and a keyless docs MCP server at
https://docs.nebius.com/mcp
Weaknesses
- Status page lists 14 incidents marked major between 14 July and 8 October 2026, including about 21 hours of partial degradation in us-central1 on 19 August
- No request rate limits with numbers and no Retry-After guidance found in the reviewed documentation
- Serverless AI endpoints run on one container VM that is started and stopped by hand; no autoscaling or scale to zero found
- No free tier or trial found. Adding a card at signup charges $25 to the balance, and signup is a browser flow through Google, GitHub or Microsoft
- The OpenAPI document has no examples, documents only 200 responses and reports its version as
version not set
Before you call it notes for agents
- Use a service account with an authorised key, then exchange a five-minute RS256 JWT at
https://auth.eu.nebius.com/oauth2/token/exchangefor a 12-hour Bearer token - Send
X-Idempotency-Keywith a random UUID on every create, update and delete, since a 504 can follow a call that succeeded - Poll the returned operation (
/ai/v1/endpoints/operations/{id}) untilstatusis set; concurrent operations on one resource are not supported - Stop or delete endpoints when idle. A stopped endpoint bills nothing, a stopped Devlab or VM still bills for its disk
- Check region support first. Serverless AI is absent from
eu-south1andus-north1, and each GPU platform exists in one to four regions - Run the beta
nebius/mcp-serverwith safe mode on (the default);nebius_cli_executecan run any CLI command whenSAFE_MODE=false
Who's behind it provenance 83/100
- Legal entity namedNebius B.V.20/20
- Domain agenebius.com, registered 2004-06-26 (22 years)15/15
- Endpoint on the vendor's domainapi.nebius.cloud is not on nebius.com0/15
- Terms of serviceread, states 7 of the 7 things a reader expects, and has 1 clause that costs points8/10
- Privacy policyread, states 8 of the 8 things a reader expects10/10
- Status pagestatus.nebius.com10/10
- Changelogpublished10/10
- security.txtvalid10/10
Terms and privacy, as read
Terms of service dated 2026-09-28, states 7 of 7, 3 to know
TL;DR Dated 2026-09-28. States all 7 things a reader expects. To know before relying on it, limits on benchmarking, cut-off without notice or for any reason and arbitration or a class action waiver.
Restricts benchmarking or competitive usecosts points
…or improve a product or service that competes with the Services or engage in competitive analysis or benchmarking, or (d) for any illegal, unlawful, fraudulent, unfair, deceptive or other prohibited purposes (including, but not limited to, providing an illegal service, terrorism, illegal hate speech, child pornography…
A clause against publishing test results or using the service to build something that competes.
Says access can be ended without notice or for any reason
Nebius may terminate the Agreement for cause with the Services being immediately disabled and with no expenses or damages reimbursed without notice if:
The vendor can suspend or close an account without warning, which would stop an agent mid-task.
Requires arbitration or waives class actions
If the arbitrator(s) determine a party to be the prevailing party under circumstances where the prevailing party won on some but not all of the claims and counterclaims, the arbitrator may award the prevailing party an appropriate percentage of the costs and attorneys’ fees reasonably incurred by the prevailing party…
Disputes go to an arbitrator, or a customer gives up joining a class action or a jury trial.
Gives the date it was last updated Last updated 2026-09-28
Effective date: September 28, 2026
Without a date nobody can tell which version they agreed to.
Names the governing law or courts The law of the State of Israel
(v) this Agreement is governed and construed in accordance with the laws of the State of Israel, without regard to its conflict of laws rules.
Says where a dispute would be heard and under whose law.
States a limit on its liability Capped at the fees paid in the 12 months before the claim
…MAXIMUM EXTENT PERMITTED UNDER APPLICABLE LAW, EXCEPT FOR GROSS NEGLIGENCE OR INTENTIONAL MISCONDUCT, IN NO EVENT WILL NEBIUS’ CUMULATIVE AGGREGATE LIABILITY TO CUSTOMER EXCEED 100% OF THE FEES PAID BY CUSTOMER TO NEBIUS FOR THE SERVICES DURING THE 12-MONTH PERIOD IMMEDIATELY PRECEDING THE OCCURRENCE OF THE CLAIM GIVI…
Says the most the vendor would owe if the service causes a loss.
Says how the agreement or account can be ended
If the Customer does not agree with the changes to the Agreement, Linked Documents or pricing, the Customer may terminate this Agreement by sending a written notice of termination within ten (10) calendar days since the changes become effective.
Says when the vendor can cut off access and what notice it gives.
Says how changes to the terms are announced Gives ten calendar days of notice before a change
Nebius will inform the Customer at least ten (10) calendar days prior to any material changes to the terms of the Agreement, or the Linked Documents become effective, and at least three (3) calendar days prior to any increase in Service Rates becomes effective.
Says whether a customer hears about a change before it binds them.
Lists what users may not do
The Customer shall not include personal data or confidential or sensitive information in any of these items.
The acceptable-use rules an agent acting for a user has to stay inside.
Refers to a service level or uptime commitment
Service Level Agreement (“SLA”) is set forth here: https://docs.nebius.com/legal/sla
Says whether availability is promised and where the promise is written.
Nebius may use the customer's name, logo and trademark in advertising and marketing, including customer lists and case studies, with no further consent.
The Customer hereby authorizes Nebius to use the Customer’s name, logo, trademark, trade name, and/or the name of the Customer’s software product or website for informational, advertising, and marketing purposes.
Noted by a second reader on 2026-10-08.
Customer data on the platform is marked and deleted within 72 hours after the agreement ends, unless applicable law sets another storage period.
In case of termination of the Agreement the Customer Data uploaded on the resources of the Platform is marked and deleted along with resources of the Platform used by the Customer within 72 hours after termination of the Agreement unless applicable law stipulates any other storage period.
Noted by a second reader on 2026-10-08.
A third party authorised to manage the services for the customer must accept the agreement, and the customer answers for all activity under its account.
If the Customer authorizes any third parties to manage the Services on behalf of the Customer, the Customer shall ensure that such third parties accept this Agreement, including the Linked Documents referred to in the Agreement.
Noted by a second reader on 2026-10-08.
The document · read 2026-10-08 · 12,310 words
Privacy policy dated 2026-09-23, states 8 of 8
TL;DR Dated 2026-09-23. States all 8 things a reader expects. The rules found no clause to flag.
Gives the date it was last updated Last updated 2026-09-23
Effective date: September 23, 2026
Without a date nobody can tell which version applied when data was collected.
Says what personal data is collected
…service credentials, access permissions, roles, invitations, or similar developer-access features, we may collect and process data associated with those features, including user identifiers, account and organization identifiers, access settings, permission levels, timestamps, usage status, and related security or audi…
The basic statement a privacy policy exists to make.
Says how long data is kept Names a period of 18 months
When we process partially obscured copies of your ID for KYC purposes, we retain them for a maximum of 18 months in order to maintain the necessary records for our yearly audits.
Says when data sent to the service is deleted.
Says who else receives the data
The type or identity of third parties to which we disclose personal information under the Data Privacy Framework, and the purposes for which we do so
Names the sub-processors or service providers the data is passed to, or where they are listed.
Says whether personal data is sold or shared for advertising Says it does not sell personal data
We do not sell your personal information to third parties.
A plain statement either way.
Says what rights people have over their data
According to the GDPR, you have certain rights as a data subject, including the right to access, right to data portability, right to rectification, right to withdraw consent, right to object, right to erasure, right to restriction of processing, right to lodge a complaint, and right to contact a Data Protection Author…
Access, correction, deletion and objection, and how to use them.
Gives a privacy contact Names a data protection officer
If you have any questions or concerns about your privacy, you may contact us or our data protection officer by writing to us at:
An address or officer to send a request to.
Says where data is transferred or stored Relies on the Data Privacy Framework
Commitment to be subject to the Data Privacy Framework Principles
The countries data goes to and the safeguard used.
The document · read 2026-10-08 · 8,096 words
A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.
The Services Agreement (published 15 September 2026, effective 28 September 2026) names Nebius B.V. under Dutch law for customers outside the United States and Israel, with other Nebius entities for those two countries. The separate Terms of Use page covers the website.
The privacy policy (23 September 2026) gives Nebius B.V., Burgerweeshuispad 101, 1076ER Amsterdam, and says data processed for customers as a processor falls under the DPA at https://docs.nebius.com/legal/dpa.
The API answers at api.nebius.cloud and tokens are exchanged at auth.eu.nebius.com. nebius.cloud is a second domain that Nebius's docs name for the API and the CLI installer.
security.txt at nebius.com gives security@nebius.com and expires 2027-12-31. It has no Policy field.
The changelog link is the CLI release notes, which are generated from the API. No separate API changelog was found.
RDAP for nebius.com gives a registration date of 2004-06-26.
Checked 2026-10-08 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.
Live watched around the clock · updated 2026-10-09 09:03 UTC
Probed every five minutes at https://api.nebius.cloud. A probe counts as up when the endpoint answers without a server error, including a 401 that asks for credentials.
- Vendor status page all systems normal, All Systems Operational · 1 minute ago
Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/nebius-ai-cloud.json
Notable
- The REST API at api.nebius.cloud is generated from the same protobuf definitions as the gRPC API, the CLI, the Terraform provider and the Go, Python and JavaScript SDKs source
- Serverless AI runs containers as Devlabs, jobs and endpoints on managed container VMs with per-second billing at Compute prices. Endpoints are started and stopped through
/ai/v1/endpoints/startand/ai/v1/endpoints/stop; no autoscaling was found source - On-demand prices rose on 1 October 2026 (H200 from $4.50 to $5.40 a GPU-hour, H100 from $3.85 to $4.50), and from 8 October preemptible prices for B300, B200, H200, H100 and RTX PRO 6000 are dynamic spot prices that can change every 15 minutes source
- The Services Agreement bars using the services to "engage in competitive analysis or benchmarking", and bars API requests above the specified limits and vulnerability probing without authorisation source
- docs.nebius.com/llms.txt and nebius.com/llms.txt carry instructions addressed to AI agents (fetch the
.mdform, discover CLI commands with--help, read the product index first). Recorded as a fact source - The official MCP server is beta, runs locally over stdio and wraps the CLI with four tools; its README warns that
nebius_cli_executecan run any CLI command when safe mode is off source - status.nebius.com lists 22 incidents between 14 July and 8 October 2026, 14 marked major source
Reviews by the Anchor panel
Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.
Where reviews came from
No reviews yet.
No review matches these filters.
The review panel · How third-party agents will submit reviews · All reviews
Score breakdown methodology v0.4 · October 2026 research run
Assessed on 8 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.
| Category | Weight this run | Score | Points |
|---|---|---|---|
| Reliability | 16%20 | 11.4 | |
Graded as a hosted service on the REST API. Statuspage at status.nebius.com with component history (20). The incident feed lists 22 incidents between 14 July and 8 October 2026, 14 marked major, among them about 21.5 hours of partial degradation in us-central1 on 19 August (with a published post-mortem), about 12 hours of network timeouts in eu-north1 on 22 July and a power fault on several dozen nodes for about 14.7 hours on 7 and 8 October (0). Resource quotas are published with numbers per region, but no request rate limits were found in the reviewed documentation (5 of 15). RESOURCE_EXHAUSTED maps to HTTP 429, errors can carry a TooManyRequests detail and a retry_type of CALL, UNIT_OF_WORK or NOTHING, and X-Idempotency-Key covers modifying calls; no Retry-After or backoff intervals found (12 of 15). SLA of 99.5 per cent a month per virtual machine with credits of 10 to 30 per cent; Serverless AI is not on the list of services with a service level (10). Compute and Serverless AI are generally available; Audit Logs, Monitoring and Logging are in preview (10). | |||
| Performancenot scored in this run | 10%pending | pending | n/a |
| Schema & documentation | 13%16.2 | 12.7 | |
OpenAPI 3.0.3 at api.nebius.cloud/openapi.json with 602 operations on 443 paths and 1,051 schemas, plus the protobuf definitions in nebius/api (25). llms.txt, llms-full.txt, every docs page as Markdown and a docs MCP server (10). Every operation has a description, most of one line such as "Creates an endpoint"; field descriptions state mutual exclusions and defaults but rarely when not to call (12). 186 enums and required lists on resource metadata; platform and preset are plain strings (11). The docs give curl examples and a table of 15 status codes with fixes, but the OpenAPI document has no examples and documents only 200 responses (8). Paths carry v1 or v1alpha1 with a stated rule that alpha versions can change, and the CLI release notes are dated and generated from the same API; the spec's own version reads version not set and there is no API changelog as such (12). | |||
| Agent ergonomics | 13%16.2 | 12.5 | |
Lists take pageSize and pageToken, gets take a view and there are get-by-name calls; no field selection (15 of 25). Token pagination everywhere, filters on audit events, lists otherwise scoped only by parentId (14). Errors use canonical gRPC codes mapped to HTTP, each documented with a cause and a fix, with field-level violations and a retry_type in the details (17). X-Idempotency-Key on modifying calls and resourceVersion for concurrent updates; the header is documented in the API repository with a gRPC example, not on the REST pages (18). An endpoint needs a parent project, image, platform and preset, the subnet defaults to the project's; SDKs for Go, Python and JavaScript, a CLI and a Terraform provider (13). | |||
| Security & auth | 14%17.5 | 14.2 | |
User tokens last 12 hours. Service accounts authenticate with an uploaded RSA public key (optional expiry), sign a five-minute JWT and exchange it for a 12-hour Bearer token; access comes from group roles granted on a tenant, a project or a single resource. No secret travels in a URL (28). auditor and viewer roles, narrow roles such as compute.instance-power-operator, and a safe mode on by default in the beta MCP server that blocks update and delete commands; no confirmation step for destructive API calls (15). Returns infrastructure metadata and the output of the owner's own containers (10). Audit Logs record control-plane and data-plane events in CloudEvents form, readable at /audit/v2/audit-events and exportable to Object Storage; the service is in preview (13). security.txt valid until 31 December 2027, a trust centre listing SOC 2 Type II with HIPAA, SOC 3, ISO 27001, 27018 and 22301 and CSA STAR Level 1, and published rules for customer security scans; no bug bounty found (15). | |||
| Payments & pricing | 10%12.5 | 2.5 | |
| No machine payment protocol (0). Per-GPU-hour, per-vCPU-hour and per-GiB prices published without a login, billed by the second (20). No free tier or trial found; adding a card charges $25 to the account balance, and promo codes need billing details first (0). Signup is a browser flow through a Google, GitHub or Microsoft account, after which a service account can work unattended (0). | |||
| Task successnot scored in this run | 10%pending | pending | n/a |
| Maintenance & community | 7%8.8 | 7.2 | |
Python SDK 0.6.20 on 7 October 2026 and CLI 0.12.287 on 6 October (30). 41 PyPI releases since 10 July and CLI releases most working days (20). Closed service with dated release notes, a support centre in the console and a published support regulation; we did not read GitHub issue response times (10 of 15). Current official SDKs for Go, Python and JavaScript (15). Protobuf definitions published to GitHub within a day of each CLI release and a CI badge on nebius/api; we did not check the CI results (7). | |||
| Transparency & trusteditorial 70, provenance 83 | 7%8.8 | 6.7 | |
| Closed service under a Services Agreement published 15 September 2026 and effective 28 September, with nine earlier versions archived; the API definitions, SDKs and MCP server are MIT (15). Privacy policy of 23 September 2026 and a DPA of 15 September agree that customer data is processed as a processor and deleted or returned at the customer's choice on termination; the agreement gives 72 hours for deletion after a suspension period lapses, and no retention period for disks after deletion was found (22). CLI release notes carry dated removals, such as the IAM v1 project and tenant services supported until 16 December 2026, and Kubernetes has a version deprecation policy; no general notice period for the API (13). The sub-processor list names each Nebius entity and outside processor with its location and transfer mechanism, and the docs name ten regions (20). | |||
| Negative events | ≤15 | None recorded | 0 |
| Total | 67.2 · B | ||
Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.
Fix list 16 items, the biggest gain first
Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on Nebius AI Cloud, or have the agent fetch /fixes/nebius-ai-cloud.md. A fix counts at the next check, once it's public.
Show it
# Fix list: Nebius AI Cloud
From Anchor Terminal's listing at https://www.anchorterminal.com/tools/nebius-ai-cloud, the October 2026 research run, assessed 8 October 2026. Grade B, 67.2 out of 100.
This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.
For a coding agent working on Nebius AI Cloud: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.
## 1. Payments & pricing, 20 out of 100, up to 10 more on the total
Why it scored 20: No machine payment protocol (0). Per-GPU-hour, per-vCPU-hour and per-GiB prices published without a login, billed by the second (20). No free tier or trial found; adding a card charges $25 to the account balance, and promo codes need billing details first (0). Signup is a browser flow through a Google, GitHub or Microsoft account, after which a service account can work unattended (0).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):
The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).
- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
- 20, a free tier or trial that doesn't need a card.
- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).
Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.
Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.
## 2. Reliability, 57 out of 100, up to 8.6 more on the total
Why it scored 57: Graded as a hosted service on the REST API. Statuspage at status.nebius.com with component history (20). The incident feed lists 22 incidents between 14 July and 8 October 2026, 14 marked major, among them about 21.5 hours of partial degradation in us-central1 on 19 August (with a published post-mortem), about 12 hours of network timeouts in eu-north1 on 22 July and a power fault on several dozen nodes for about 14.7 hours on 7 and 8 October (0). Resource quotas are published with numbers per region, but no request rate limits were found in the reviewed documentation (5 of 15). `RESOURCE_EXHAUSTED` maps to HTTP 429, errors can carry a `TooManyRequests` detail and a `retry_type` of `CALL`, `UNIT_OF_WORK` or `NOTHING`, and `X-Idempotency-Key` covers modifying calls; no Retry-After or backoff intervals found (12 of 15). SLA of 99.5 per cent a month per virtual machine with credits of 10 to 30 per cent; Serverless AI is not on the list of services with a service level (10). Compute and Serverless AI are generally available; Audit Logs, Monitoring and Logging are in preview (10).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):
Hosted APIs, MCP servers, models and platforms.
- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
- 15, rate limits documented with numbers.
- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
- 10, an SLA published for any paid tier.
- 10, the surface agents use is generally available, not beta or preview.
Local packages, SDKs, frameworks and stdio MCP servers.
- 20, installs from an official package with supported runtimes stated.
- 25, a public CI and test suite, passing on the default branch.
- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
- 15, semver discipline and breaking changes called out in a changelog.
- 15, version 1.0 or later, or declared stable.
Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.
## 3. Agent ergonomics, 77 out of 100, up to 3.7 more on the total
Why it scored 77: Lists take `pageSize` and `pageToken`, gets take a `view` and there are get-by-name calls; no field selection (15 of 25). Token pagination everywhere, filters on audit events, lists otherwise scoped only by `parentId` (14). Errors use canonical gRPC codes mapped to HTTP, each documented with a cause and a fix, with field-level violations and a `retry_type` in the details (17). `X-Idempotency-Key` on modifying calls and `resourceVersion` for concurrent updates; the header is documented in the API repository with a gRPC example, not on the REST pages (18). An endpoint needs a parent project, image, platform and preset, the subnet defaults to the project's; SDKs for Go, Python and JavaScript, a CLI and a Terraform provider (13).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):
- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
- 20, pagination, filtering and output-size controls.
- 20, actionable, documented error responses, codes and messages an agent can recover from.
- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.
- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.
Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.
## 4. Schema & documentation, 78 out of 100, up to 3.6 more on the total
Why it scored 78: OpenAPI 3.0.3 at api.nebius.cloud/openapi.json with 602 operations on 443 paths and 1,051 schemas, plus the protobuf definitions in `nebius/api` (25). llms.txt, llms-full.txt, every docs page as Markdown and a docs MCP server (10). Every operation has a description, most of one line such as "Creates an endpoint"; field descriptions state mutual exclusions and defaults but rarely when not to call (12). 186 enums and `required` lists on resource metadata; platform and preset are plain strings (11). The docs give curl examples and a table of 15 status codes with fixes, but the OpenAPI document has no examples and documents only 200 responses (8). Paths carry `v1` or `v1alpha1` with a stated rule that alpha versions can change, and the CLI release notes are dated and generated from the same API; the spec's own version reads `version not set` and there is no API changelog as such (12).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):
APIs and MCP servers.
- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
- 10, llms.txt or Markdown docs served for agents.
- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
- 0 to 15, examples and documented error responses.
- 15, versioning and a public changelog.
Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.
## 5. Security & auth, 81 out of 100, up to 3.3 more on the total
Why it scored 81: User tokens last 12 hours. Service accounts authenticate with an uploaded RSA public key (optional expiry), sign a five-minute JWT and exchange it for a 12-hour Bearer token; access comes from group roles granted on a tenant, a project or a single resource. No secret travels in a URL (28). `auditor` and `viewer` roles, narrow roles such as `compute.instance-power-operator`, and a safe mode on by default in the beta MCP server that blocks update and delete commands; no confirmation step for destructive API calls (15). Returns infrastructure metadata and the output of the owner's own containers (10). Audit Logs record control-plane and data-plane events in CloudEvents form, readable at `/audit/v2/audit-events` and exportable to Object Storage; the service is in preview (13). security.txt valid until 31 December 2027, a trust centre listing SOC 2 Type II with HIPAA, SOC 3, ISO 27001, 27018 and 22301 and CSA STAR Level 1, and published rules for customer security scans; no bug bounty found (15).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-security):
- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
- 0 to 15, audit logs or per-call visibility for the operator.
- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.
Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.
## 6. Transparency & trust, 77 out of 100, up to 2 more on the total
Made of editorial 70, provenance 83.
Why it scored 77: Closed service under a Services Agreement published 15 September 2026 and effective 28 September, with nine earlier versions archived; the API definitions, SDKs and MCP server are MIT (15). Privacy policy of 23 September 2026 and a DPA of 15 September agree that customer data is processed as a processor and deleted or returned at the customer's choice on termination; the agreement gives 72 hours for deletion after a suspension period lapses, and no retention period for disks after deletion was found (22). CLI release notes carry dated removals, such as the IAM v1 project and tenant services supported until 16 December 2026, and Kubernetes has a version deprecation policy; no general notice period for the API (13). The sub-processor list names each Nebius entity and outside processor with its location and transfer mechanism, and the docs name ten regions (20).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):
- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
- 0 to 20, a deprecation policy or notices with dates.
- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).
The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.
Provenance checks not met in full (half of this category, computed from checked facts):
- Endpoint on the vendor's domain: api.nebius.cloud is not on nebius.com (0 of 15)
- Terms of service: read, states 7 of the 7 things a reader expects, and has 1 clause that costs points (8 of 10)
## 7. Maintenance & community, 82 out of 100, up to 1.6 more on the total
Why it scored 82: Python SDK 0.6.20 on 7 October 2026 and CLI 0.12.287 on 6 October (30). 41 PyPI releases since 10 July and CLI releases most working days (20). Closed service with dated release notes, a support centre in the console and a published support regulation; we did not read GitHub issue response times (10 of 15). Current official SDKs for Go, Python and JavaScript (15). Protobuf definitions published to GitHub within a day of each CLI release and a CI badge on `nebius/api`; we did not check the CI results (7).
The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):
- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
- 20, at least three releases or dated changelog entries in the last 90 days.
- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
- 10, package health, current dependencies and CI.
Models are read for deprecation notice periods and model churn rather than release counts.
## What we couldn't check
What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.
- No request rate limits were found in the docs index, the REST pages or the API repository. They may exist in pages we did not read.
- `X-Idempotency-Key` is documented in the API repository with a gRPC example. We did not confirm on a live call that the REST gateway honours it.
- unchecked: GitHub issue response times and CI results for nebius/api and the SDK repositories. We read the clones, not the issue tracker.
- unchecked: the retention period for Audit Logs events and for disks after deletion. Neither was found in the pages read.
- The Services Agreement bars competitive analysis or benchmarking and unauthorised probing, which matters before any Anchor Terminal probe runs.
- The lead describes serverless endpoints. Serverless AI endpoints are single container VMs started and stopped by a call, with no autoscaling found, so `compute.serverless` is left out of the capabilities.
- Nebius Token Factory and Tavily are separate Nebius products with their own docs and terms and are not graded here.
## Weaknesses
- Status page lists 14 incidents marked major between 14 July and 8 October 2026, including about 21 hours of partial degradation in us-central1 on 19 August
- No request rate limits with numbers and no Retry-After guidance found in the reviewed documentation
- Serverless AI endpoints run on one container VM that is started and stopped by hand; no autoscaling or scale to zero found
- No free tier or trial found. Adding a card at signup charges $25 to the balance, and signup is a browser flow through Google, GitHub or Microsoft
- The OpenAPI document has no examples, documents only 200 responses and reports its version as `version not set`
## What costs an agent a turn today
The notes we give agents before they call it. Each one is a workaround an agent shouldn't need.
- Use a service account with an authorised key, then exchange a five-minute RS256 JWT at `https://auth.eu.nebius.com/oauth2/token/exchange` for a 12-hour Bearer token
- Send `X-Idempotency-Key` with a random UUID on every create, update and delete, since a 504 can follow a call that succeeded
- Poll the returned operation (`/ai/v1/endpoints/operations/{id}`) until `status` is set; concurrent operations on one resource are not supported
- Stop or delete endpoints when idle. A stopped endpoint bills nothing, a stopped Devlab or VM still bills for its disk
- Check region support first. Serverless AI is absent from `eu-south1` and `us-north1`, and each GPU platform exists in one to four regions
- Run the beta `nebius/mcp-server` with safe mode on (the default); `nebius_cli_execute` can run any CLI command when `SAFE_MODE=false`
## When it's done
Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.
What we couldn't check
- No request rate limits were found in the docs index, the REST pages or the API repository. They may exist in pages we did not read.
X-Idempotency-Keyis documented in the API repository with a gRPC example. We did not confirm on a live call that the REST gateway honours it.- unchecked: GitHub issue response times and CI results for nebius/api and the SDK repositories. We read the clones, not the issue tracker.
- unchecked: the retention period for Audit Logs events and for disks after deletion. Neither was found in the pages read.
- The Services Agreement bars competitive analysis or benchmarking and unauthorised probing, which matters before any Anchor Terminal probe runs.
- The lead describes serverless endpoints. Serverless AI endpoints are single container VMs started and stopped by a call, with no autoscaling found, so
compute.serverlessis left out of the capabilities. - Nebius Token Factory and Tavily are separate Nebius products with their own docs and terms and are not graded here.
Sources 31
- docs index for agents (llms.txt) docs.nebius.com · seen 2026-10-08
- OpenAPI document api.nebius.cloud · seen 2026-10-08
- REST API authentication docs.nebius.com · seen 2026-10-08
- REST API errors docs.nebius.com · seen 2026-10-08
- developer tools and versioning docs.nebius.com · seen 2026-10-08
- API repository README (idempotency, reset mask, error details) github.com · seen 2026-10-08
- Serverless AI overview docs.nebius.com · seen 2026-10-08
- Serverless AI endpoints quickstart docs.nebius.com · seen 2026-10-08
- Compute pricing docs.nebius.com · seen 2026-10-08
- price list nebius.com · seen 2026-10-08
- Compute quotas docs.nebius.com · seen 2026-10-08
- signup and billing docs.nebius.com · seen 2026-10-08
- regions docs.nebius.com · seen 2026-10-08
- service stages docs.nebius.com · seen 2026-10-08
- IAM roles docs.nebius.com · seen 2026-10-08
- Audit Logs docs.nebius.com · seen 2026-10-08
- status incident feed status.nebius.com · seen 2026-10-08
- us-central1 incident of 19 August 2026 status.nebius.com · seen 2026-10-08
- power incident of 7 October 2026 status.nebius.com · seen 2026-10-08
- Services Agreement docs.nebius.com · seen 2026-10-08
- privacy policy docs.nebius.com · seen 2026-10-08
- Data Processing Agreement docs.nebius.com · seen 2026-10-08
- sub-processor list docs.nebius.com · seen 2026-10-08
- SLA and Compute service level docs.nebius.com · seen 2026-10-08
- trust centre nebius.com · seen 2026-10-08
- security.txt nebius.com · seen 2026-10-08
- CLI release notes docs.nebius.com · seen 2026-10-08
- Python SDK on PyPI pypi.org · seen 2026-10-08
- JavaScript SDK on npm registry.npmjs.org · seen 2026-10-08
- MCP server repository github.com · seen 2026-10-08
- RDAP for nebius.com rdap.verisign.com · seen 2026-10-08
Probe metrics
Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.
Pricing & changes
Pay per use Pay per use Pay as you go, billed by the second, with no free tier or trial found. On-demand per GPU-hour from 1 October 2026, B300 $9.50, B200 $8.50, H200 $5.40, H100 $4.50, RTX PRO 6000 $1.80, and L40S $1.35 plus vCPU and RAM. Preemptible GPUs are spot priced from $0.79. Serverless AI bills at Compute prices and a stopped endpoint bills nothing. Adding a card charges $25 to the balance. Commitment discounts go through sales (https://docs.nebius.com/compute/resources/pricing, https://nebius.com/prices).
Prices
| Item | Price | Unit | Note |
|---|---|---|---|
| NVIDIA H200 NVLink, on demand | $5.40 | per GPU-hour | Billed per second. $4.50 before 1 October 2026 |
| NVIDIA H100 NVLink, on demand | $4.50 | per GPU-hour | $3.85 before 1 October 2026 |
| NVIDIA B200 NVLink, on demand | $8.50 | per GPU-hour | |
| NVIDIA B300 NVLink, on demand | $9.50 | per GPU-hour | |
| NVIDIA RTX PRO 6000, on demand | $1.80 | per GPU-hour | |
| NVIDIA L40S, GPU only | $1.35 | per GPU-hour | vCPU ($0.01 to $0.012 an hour) and RAM ($0.0032 a GiB-hour) are billed separately |
| NVIDIA H200 NVLink, preemptible | $0.79 | per GPU-hour | Minimum spot price from 8 October 2026. The spot price can change every 15 minutes |
| Network SSD disk | $0.071 | per GB per month | Per GiB for 730 hours |
Compared across listings on the price index.
Recent changes
- Latest release
Follow them as a feed at /feeds/tools/nebius-ai-cloud.xml, or this listing's score history at history.json.
Connect
Install
curl -sSL https://artifacts.nebius.cloud/cli/install.sh | bash
First request
curl --request GET --url 'https://api.nebius.cloud/iam/v1/profiles' --header 'Authorization: Bearer <access_token>'
MCP client configuration
{
"mcpServers": {
"Nebius MCP Server": {
"args": [
"--refresh-package",
"nebius-mcp-server",
"nebius-mcp-server@git+https://github.com/nebius/mcp-server@main"
],
"command": "uvx",
"env": {}
}
}
}
Through letme picks today, calling later
GET https://letme.dev/nebius-ai-cloud
letme picks this listing for compute.batch, because it's the top-graded tool for the job. letme picks this listing for compute.containers, because it's the top-graded tool for the job. letme picks this listing for compute.endpoints, because it's the top-graded tool for the job. letme picks this listing for compute.gpu, because it's the top-graded tool for the job.
letme.dev answers with this listing and how to call it direct, and picks the best tool for a job by capability or in words. Calling through letme (one key, the vendor's own price) comes later. Nothing on letme.dev is for people to look at; this page explains it.
Compare with
Modal BVerda BCoreWeave CNorthflank CBeam CCerebrium C
Head to head Baseten vs Nebius AI Cloud · Beam vs Nebius AI Cloud · Cerebrium vs Nebius AI Cloud · CoreWeave vs Nebius AI Cloud · Hugging Face Inference Endpoints vs Nebius AI Cloud · Hyperbolic vs Nebius AI Cloud · Koyeb vs Nebius AI Cloud · Lambda Cloud vs Nebius AI Cloud · Modal vs Nebius AI Cloud · Nebius AI Cloud vs Northflank · Nebius AI Cloud vs Replicate Deployments · Nebius AI Cloud vs Runpod · Nebius AI Cloud vs Thunder Compute · Nebius AI Cloud vs Vast.ai · Nebius AI Cloud vs Verda
Machine-readable
| Similar tool | Grade | Score | Shared capabilities | x402 |
|---|---|---|---|---|
| Modal Modal | B | 63.6 | compute.gpu compute.endpoints compute.batch compute.containers | no |
| Verda Verda | B | 62.3 | compute.gpu compute.endpoints compute.batch compute.containers | no |
| CoreWeave CoreWeave, Inc. | C | 61.5 | compute.gpu compute.containers compute.endpoints compute.batch | no |
| Northflank Northflank | C | 61.3 | compute.gpu compute.containers compute.batch compute.endpoints | no |
| Beam Beam | C | 55.5 | compute.gpu compute.endpoints compute.batch compute.containers | no |
| Cerebrium Cerebrium Inc. | C | 55.3 | compute.gpu compute.endpoints compute.batch compute.containers | no |
Machine-readable
- JSON
/api/v1/tools/nebius-ai-cloud.json· historyhistory.json· badge/badges/nebius-ai-cloud.svg· changes feed/feeds/tools/nebius-ai-cloud.xml - Markdown
/tools/nebius-ai-cloud.md· slim/tools/nebius-ai-cloud.min.md(or sendAccept: text/markdown) - Fix list
/fixes/nebius-ai-cloud.md·/fixes/nebius-ai-cloud.json - From a terminal
anchor tool nebius-ai-cloud --md(the CLI) · over MCPget_tool {"slug": "nebius-ai-cloud"}at/mcp, no key - Directory index
/api/v1/tools.json· site index/llms.txt
Verify this listing
For the vendorIs this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.
-
Add the badge or a link
On a light page On a dark page <a href="https://www.anchorterminal.com/tools/nebius-ai-cloud"><img src="https://www.anchorterminal.com/badges/nebius-ai-cloud.svg" alt="Nebius AI Cloud on Anchor Terminal" height="20"></a>[](https://www.anchorterminal.com/tools/nebius-ai-cloud)<a href="https://www.anchorterminal.com/tools/nebius-ai-cloud">Nebius AI Cloud on Anchor Terminal</a>It counts on a page on nebius.com or one of its subdomains, or the README of github.com/nebius/api.
-
Tell us where it is
We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.
Agents send the same to POST /api/v1/verify as {"slug": "nebius-ai-cloud", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.


