Best of · Models & inference
Best GPU and serverless compute for AI workloads
The 10 highest-scoring of 19 GPU and serverless compute for AI workloads on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 19 ranked
- 1 agent-ready
- 17 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Schema & documentation
85/100 on schema & documentation, against 79 for the overall leader.
Security & auth
Hugging Face Inference Endpoints B
83/100 on security & auth, against 76 for the overall leader.
Transparency & trust
78/100 on transparency & trust, against 72 for the overall leader.
Lowest paid price per vCPU hour
$0.0167 per vCPU hour, the lowest of the 7 listings here with a paid price in this unit (free allowances aside).
Also Cerebrium, $0.0236 per vCPU hour.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Modal Sandboxes Modal |
BB 75.5 | GPU work inside a sandbox, or agents already running on Modal. | $0.071 / vCPU-hr | local |
| 2 | Nebius AI Cloud Nebius |
B 67.2 | Teams that want whole GPU VMs or InfiniBand clusters in Europe, the UK, Israel or the US with IAM, Terraform and an SLA, and are content to manage endpoint lifecycles themselves. | Pay per use | hosted and local |
| 3 | Baseten Baseten |
B 66.5 | Teams that want one model behind a production endpoint with real autoscaling knobs, environments and scoped keys. | Pay per use | hosted |
| 4 | Hugging Face Inference Endpoints Hugging Face, Inc. |
B 64.5 | Teams whose models already live on the Hugging Face Hub and who want a dedicated endpoint on a named cloud and region with standard open-source engines. | $0.033 / vCPU-hr | hosted |
| 5 | Modal Modal |
B 63.6 | Python teams that want GPU functions, batch jobs and HTTP endpoints from one decorator with scale to zero. | $250 / mo | local |
| 6 | Replicate Deployments Replicate |
B 63.6 | Teams already calling Replicate's public models who want their own model behind the same API, MCP server and webhooks. | Pay per use | hosted and local |
| 7 | Crusoe Cloud Crusoe |
B 63.4 | Teams that want whole GPU machines or clusters with managed Kubernetes or Slurm in the US, Iceland or Norway, under an SLA, and are content to manage lifecycles themselves. | $0.04 / vCPU-hr | hosted and local |
| 8 | Vast.ai Vast.ai Inc. |
B 62.6 | Cost-sensitive training, batch work and self-managed inference where the agent can search listings, set a price cap and tolerate host variance or interruption. | Pay per use | hosted |
| 9 | Verda Verda |
B 62.3 | Agents that rent whole GPU machines or clusters in Finland for training or batch work, or deploy scale-to-zero container endpoints, and that can run a local CLI for MCP. | Pay per use | hosted and local |
| 10 | CoreWeave CoreWeave, Inc. |
C 61.5 | Teams that already run Kubernetes and need whole multi-GPU nodes or clusters for training and dedicated inference, with a contract and IAM. | Pay per use | hosted |
9 more are ranked in the full table.
How to choose
- Cold start from zeroCheck the cold start time from zero and whether it includes loading model weights, because an agent scaled to zero waits through it before answering.
- Billing per second or hourCheck whether billing is per second or per hour, and whether idle warm instances are charged, since that sets the real cost of bursty agent traffic.
- GPU types and memoryCheck which GPU types each region has and how much memory each type holds, because a model that does not fit in memory cannot run.
- Autoscaling and scale to zeroCheck how quickly the endpoint scales up under load and whether it can scale to zero, because a slow scale-up makes concurrent agent calls queue or time out.
How the benchmark tests this category. The same model deployed as an endpoint on each platform, called cold and warm, then scaled to zero. We time cold starts, check the scaling and add up the cost per GPU-hour.
Each one in detail
Modal Sandboxes
BB 75.5/100Modal's sandboxed compute environments for running code, with SDK access, GPU support and filesystem snapshots.
Verdict GPU sandboxes at the same per-second rates as the rest of Modal. No REST API, and the JavaScript and Go SDKs are beta.
Choose it for GPU work inside a sandbox, or agents already running on Modal.
Strengths
- GPU sandboxes at the same per-second rates as the rest of Modal
- Outbound traffic blockable or limited to CIDR ranges, and no inbound connections without tunnels
- $30 of compute every month on Starter, no card
Weaknesses
- No REST API, and the JavaScript and Go SDKs are beta
- Default lifetime of 5 minutes and a hard maximum of 24 hours
- gVisor rather than a VM unless you're on Team or Enterprise for the VM runtime
Price $0.071 / vCPU-hrAuth API keyx402 nolocal
Nebius AI Cloud
B 67.2/100Nebius AI Cloud rents NVIDIA GPU virtual machines and InfiniBand clusters, with managed Kubernetes, Slurm and Serverless AI jobs and endpoints for containers. Resources are managed through REST and gRPC APIs, a CLI, a Terraform provider and SDKs.
Verdict One API definition generates the REST and gRPC interfaces, the CLI, Terraform provider and three SDKs, with a 602-operation OpenAPI document, X-Idempotency-Key and role-scoped service accounts. The status page lists 14 major incidents between 14 July and 8 October 2026, no request rate limits were found, and signup needs a browser and a card.
Choose it for Teams that want whole GPU VMs or InfiniBand clusters in Europe, the UK, Israel or the US with IAM, Terraform and an SLA, and are content to manage endpoint lifecycles themselves.
Strengths
- OpenAPI 3.0.3 document at
https://api.nebius.cloud/openapi.jsonwith 602 operations, generated from the same protobuf definitions as the gRPC API, CLI, Terraform provider and SDKs X-Idempotency-Keyheader for modifying calls, and aretry_typefield on errors that says whether to retry the call- Service accounts sign in with an uploaded RSA key and receive 12-hour tokens, with roles granted per tenant, project or resource
Weaknesses
- Status page lists 14 incidents marked major between 14 July and 8 October 2026, including about 21 hours of partial degradation in us-central1 on 19 August
- No request rate limits with numbers and no Retry-After guidance found in the reviewed documentation
- Serverless AI endpoints run on one container VM that is started and stopped by hand; no autoscaling or scale to zero found
Price Pay per useAuth OAuth or keyx402 nohosted and local
Baseten
B 66.5/100Dedicated model deployments packaged with the open-source Truss framework and served behind a per-model HTTPS endpoint, with autoscaling from zero replicas, async inference, a management API and per-minute GPU billing from T4 to B200.
Verdict Team API keys scoped to inference-only, metrics-only or a single environment or model, plus a Viewer role since 1 September 2026. H100 at $6.50 and A100 at $4.00 an hour, and start-up and idle replica time are billed.
Choose it for Teams that want one model behind a production endpoint with real autoscaling knobs, environments and scoped keys.
Strengths
- Team API keys scoped to inference-only, metrics-only or a single environment or model, plus a Viewer role since 1 September 2026
- Public OpenAPI spec for the management API at api.baseten.co/v1/spec, and llms.txt with Markdown twins
- Rate limits published per endpoint with a
retry_afterfield on 429
Weaknesses
- H100 at $6.50 and A100 at $4.00 an hour, and start-up and idle replica time are billed
- 21 status-page incidents between 31 July and 29 September 2026, mostly single-cluster 5xx
- A bot token with admin access to the product and GitOps repositories sat exposed from March 2023 until July 2026
Price Pay per useAuth API keyx402 nohosted
Hugging Face Inference Endpoints
B 64.5/100Managed Hugging Face service that deploys a Hub model as a dedicated, autoscaling HTTPS endpoint on AWS, Azure or Google Cloud, using vLLM, TGI, SGLang, llama.cpp, TEI or a custom container. Managed by REST API, Python client, CLI or MCP.
Verdict OAuth scopes separate reading endpoints from writing them, both OpenAPI documents are public, and the unauthenticated /v2/provider route lists every instance with its hourly price. An account needs a payment method and credits before the first deployment, no rate limits or SLA were found for the management API, and the docs price table disagrees with the live list in places.
Choose it for Teams whose models already live on the Hugging Face Hub and who want a dedicated endpoint on a named cloud and region with standard open-source engines.
Strengths
- Public OpenAPI 3.1 documents for the management API (46 operations) and the catalogue API (3), plus llms.txt and a Markdown twin of every docs page
- The MCP server at endpoints.huggingface.co/mcp uses OAuth with
read-endpointsandwrite-endpointsscopes, PKCE and dynamic client registration GET /v2/providerneeds no token and returns each instance type by cloud and region with status and price per hour
Weaknesses
- No free tier. The docs require a payment method and credits, and replicas are billed while initialising as well as running
- No rate limits, SLA or idempotency keys were found for the management API, and its OpenAPI document lists only 200 responses on 43 of 46 operations
- The docs price table and the live provider list disagree. Inferentia2 x1 is $0.75 in the docs and $1.95 in the API, and AWS H200 is listed in the docs and marked deprecated in the API
Price $0.033 / vCPU-hrAuth OAuth or keyx402 nohosted
Modal
B 63.6/100Serverless functions, web endpoints, servers and GPU jobs from a Python decorator, with JavaScript and Go SDKs.
Verdict Scale to zero by default, per-second billing and about one-second container boots. No REST API or OpenAPI spec for deploying or invoking Functions.
Choose it for Python teams that want GPU functions, batch jobs and HTTP endpoints from one decorator with scale to zero.
Strengths
- Scale to zero by default, per-second billing and about one-second container boots
- Retention stated per data type (inputs and outputs up to 7 days, logs 1 to 30 days)
- Python, JavaScript and Go SDKs, with llms.txt and dated release notes
Weaknesses
- No REST API or OpenAPI spec for deploying or invoking Functions
- Web endpoints are open by default until proxy tokens are added
- RBAC, audit logs and HIPAA only on Enterprise
Price $250 / moAuth API keyx402 nolocal
Replicate Deployments
B 63.6/100Replicate's service for deploying and running custom models.
Verdict OpenAPI file, llms.txt and an MCP server with a two-tool code mode. Private instances bill set-up and idle time, H100 at $5.49 an hour.
Choose it for Teams already calling Replicate's public models who want their own model behind the same API, MCP server and webhooks.
Strengths
- OpenAPI file, llms.txt and an MCP server with a two-tool code mode
- Deployment min and max instances settable over the API, 0 allowed
- API prediction data deleted after one hour by default
Weaknesses
- Private instances bill set-up and idle time, H100 at $5.49 an hour
- API tokens have no scopes, expiry or audit log
- Changelog silent since 21 April 2026
Price Pay per useAuth API keyx402 nohosted and local
Crusoe Cloud
B 63.4/100Crusoe Cloud rents NVIDIA and AMD GPU virtual machines with managed Kubernetes and Slurm, and runs hosted model inference, dedicated deployments and LoRA fine-tuning. Resources are managed through a REST API, a CLI, a Terraform provider and a Go client.
Verdict The v1 REST API has a published OpenAPI description, Markdown docs, HMAC-signed keys, reader roles and a 90-day audit log, with on-demand GPU prices public. No idempotency keys or control-plane rate limits were found, spot and B200 prices go through sales, and a storage outage in eu-norway1 lasted about four hours on 29 September 2026.
Choose it for Teams that want whole GPU machines or clusters with managed Kubernetes or Slurm in the US, Iceland or Norway, under an SLA, and are content to manage lifecycles themselves.
Strengths
- Requests are signed with HMAC-SHA256 over the path, query, verb and timestamp, so the secret key never travels with a request.
- The API description is published for
/v1and/v1alpha5, with 1,231 examples, and the docs servellms.txtand a Markdown twin of each page. - On-demand prices are public and billed per second. H100 is $3.90 and H200 $4.29 a GPU-hour, with no charge for ingress or egress.
Weaknesses
- No idempotency key or safe-retry guidance was found for create and delete calls, which return asynchronous operations to poll.
- No request rate limits were found for the infrastructure API. Numbers are published only for Serverless Inference.
- Spot prices and on-demand prices for GB200, B200 and MI355X are listed as contact sales, and provisioning compute needs a non-prepaid credit card.
Price $0.04 / vCPU-hrAuth API keyx402 nohosted and local
Vast.ai
B 62.6/100Vast.ai is a marketplace for renting GPUs by the second from independent hosts and data centres, as Docker instances, virtual machines or autoscaling serverless endpoints. Agents use the vastai CLI, a Python SDK or a REST API.
Verdict API keys can be limited by permission category and by resource ID, the REST API has a public OpenAPI 3.1 file, and the CLI ships weekly with a skill file for coding agents. Machines belong to independent hosts, prices move with the market, rate-limit thresholds are unpublished, and there is no SLA or free tier.
Choose it for Cost-sensitive training, batch work and self-managed inference where the agent can search listings, set a price cap and tolerate host variance or interruption.
Strengths
- API keys take 11 permission categories and per-endpoint constraints on resource IDs, and can be reset or deleted at once
- Public OpenAPI 3.1 file with 89 operations, llms.txt and Markdown docs pages
- MIT CLI and Python SDK, v1.8.3 on 2 October 2026, with 12 tagged releases since 27 July 2026
Weaknesses
- No SLA. The terms say availability is not guaranteed and the service can change without notice
- Rate-limit thresholds are unpublished and 429 responses carry no
Retry-Afterheader - No free tier. Credit is prepaid with a $5 minimum deposit after a browser signup and email verification
Price Pay per useAuth API keyx402 nohosted
Verda
B 62.3/100Verda, formerly DataCrunch, is a Finnish GPU cloud renting GPU and CPU instances, clusters, serverless containers and storage. Agents use its REST API, Python and Go SDKs, or a CLI with a built-in MCP server in beta.
Verdict The REST API has a public OpenAPI 3.1 document, a dated changelog, published rate limits with Retry-After, and an audit log endpoint. The CLI's MCP server refuses billed or destructive calls without confirm: true. Access needs a browser signup and a prepaid balance, credentials carry one scope, and instance creation has no idempotency key.
Choose it for Agents that rent whole GPU machines or clusters in Finland for training or batch work, or deploy scale-to-zero container endpoints, and that can run a local CLI for MCP.
Strengths
- OpenAPI 3.1 document with 112 operations and 156 schemas, plus
llms.txtfiles on three hosts and the whole docs corpus as Markdown - Rate limits of 500 requests a minute per project and 60 per endpoint, with
RateLimitheaders andRetry-Afteron 429 - API changelog with 11 dated entries between 11 September and 7 October 2026
Weaknesses
- No free tier. Accounts are prepaid, and instances are discontinued and volumes deleted when the balance reaches zero
- Cloud API credentials carry one scope,
cloud-api-v1, with no read-only credential found in the reviewed documentation POST /v1/instanceshas no idempotency key, so a retried launch can start a second billed instance
Price Pay per useAuth OAuthx402 nohosted and local
CoreWeave
C 61.5/100CoreWeave is a GPU cloud that rents NVIDIA GPU nodes through a managed Kubernetes service (CKS), with REST and gRPC platform APIs, a Terraform provider and a hosted MCP server for observability.
Verdict Per-hour prices for 8-GPU H100, H200 and B200 nodes are public, the docs ship as Markdown with llms.txt and embedded OpenAPI specs, and IAM has read-only roles per service. An organisation must be approved by CoreWeave's sales team before any token exists, and no request rate limits were found in the reviewed documentation.
Choose it for Teams that already run Kubernetes and need whole multi-GPU nodes or clusters for training and dedicated inference, with a contract and IAM.
Strengths
- On-demand and spot prices per instance-hour are public, for example HGX H100 at $49.24 on demand and $19.71 spot for 8 GPUs
- Docs are served as Markdown with an llms.txt index, and OpenAPI 3.0 specs for CKS, VPC, storage and inference are embedded in the reference pages
- IAM has Viewer and Admin roles per service, API tokens carry an expiry, and Kubernetes audit logs can be forwarded with Telemetry Relay
Weaknesses
- No self-serve route. CoreWeave's sales team approves an organisation and emails the activation link before a token can be created
- No request rate limits, 429 guidance or idempotency keys found for api.coreweave.com in the reviewed documentation
- The CKS API is versioned v1beta1 and the Node Pool resource and Dedicated Inference API v1alpha1
Price Pay per useAuth Tokenx402 nohosted
Head to head
- Baseten vs Nebius AI Cloud B 66.5 vs B 67.2
- Hugging Face Inference Endpoints vs Nebius AI Cloud B 64.5 vs B 67.2
- Modal vs Nebius AI Cloud B 63.6 vs B 67.2
- Baseten vs Hugging Face Inference Endpoints B 66.5 vs B 64.5
- Baseten vs Modal B 66.5 vs B 63.6
- Hugging Face Inference Endpoints vs Modal B 64.5 vs B 63.6
Questions
What are the highest-rated GPU and serverless compute for AI workloads for AI agents?
Modal Sandboxes has the highest benchmark score of the 19 ranked GPU and serverless compute for AI workloads, 75.5 (BB). Nebius AI Cloud is second with 67.2 (B).
How many GPU and serverless compute for AI workloads are agent-ready?
1 of the 19 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which GPU and serverless compute for AI workloads accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
Which of these GPU and serverless compute for AI workloads is cheapest?
By published paid prices, Northflank, at $0.0167 per vCPU hour, the lowest of the 7 listings here with a paid price in this unit (free allowances aside). Plans, volume tiers and free allowances change the sum, so check the listing's price table.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 167 head-to-head comparisons · Best tools in every category