# Best GPU and serverless compute for AI workloads (slim) > Modal Sandboxes (BB), Nebius AI Cloud (B) and Baseten (B) lead the 19 ranked GPU and serverless compute for AI workloads. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/gpu-compute/index.md (~5,950 tokens) · this version ~1,530 tokens · JSON https://www.anchorterminal.com/best/gpu-compute/index.json · canonical https://www.anchorterminal.com/best/gpu-compute/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 The 10 highest-scoring of 19 GPU and serverless compute for AI workloads on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 19 · agent-ready (BB or better): 1 · accept x402: 0 · hosted endpoints: 17 - Full ranked table: https://www.anchorterminal.com/categories/gpu-compute.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/gpu-compute/index.md (167) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md) | BB | 75.5 | GPU work inside a sandbox, or agents already running on Modal. | $0.071 / vCPU-hr | local | | 2 | [Nebius AI Cloud](https://www.anchorterminal.com/tools/nebius-ai-cloud.md) | B | 67.2 | Teams that want whole GPU VMs or InfiniBand clusters in Europe, the UK, Israel or the US with IAM, Terraform and an SLA, and are content to manage endpoint lifecycles themselves. | Pay per use | hosted and local | | 3 | [Baseten](https://www.anchorterminal.com/tools/baseten.md) | B | 66.5 | Teams that want one model behind a production endpoint with real autoscaling knobs, environments and scoped keys. | Pay per use | hosted | | 4 | [Hugging Face Inference Endpoints](https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.md) | B | 64.5 | Teams whose models already live on the Hugging Face Hub and who want a dedicated endpoint on a named cloud and region with standard open-source engines. | $0.033 / vCPU-hr | hosted | | 5 | [Modal](https://www.anchorterminal.com/tools/modal.md) | B | 63.6 | Python teams that want GPU functions, batch jobs and HTTP endpoints from one decorator with scale to zero. | $250 / mo | local | | 6 | [Replicate Deployments](https://www.anchorterminal.com/tools/replicate-deploy.md) | B | 63.6 | Teams already calling Replicate's public models who want their own model behind the same API, MCP server and webhooks. | Pay per use | hosted and local | | 7 | [Crusoe Cloud](https://www.anchorterminal.com/tools/crusoe-cloud.md) | B | 63.4 | Teams that want whole GPU machines or clusters with managed Kubernetes or Slurm in the US, Iceland or Norway, under an SLA, and are content to manage lifecycles themselves. | $0.04 / vCPU-hr | hosted and local | | 8 | [Vast.ai](https://www.anchorterminal.com/tools/vast-ai.md) | B | 62.6 | Cost-sensitive training, batch work and self-managed inference where the agent can search listings, set a price cap and tolerate host variance or interruption. | Pay per use | hosted | | 9 | [Verda](https://www.anchorterminal.com/tools/verda.md) | B | 62.3 | Agents that rent whole GPU machines or clusters in Finland for training or batch work, or deploy scale-to-zero container endpoints, and that can run a local CLI for MCP. | Pay per use | hosted and local | | 10 | [CoreWeave](https://www.anchorterminal.com/tools/coreweave.md) | C | 61.5 | Teams that already run Kubernetes and need whole multi-GPU nodes or clusters for training and dedicated inference, with a contract and IAM. | Pay per use | hosted | ## Picks by need - Highest score overall: [Modal Sandboxes](https://www.anchorterminal.com/tools/modal-sandboxes.md), BB, 75.5/100 on the benchmark. Also [Nebius AI Cloud](https://www.anchorterminal.com/tools/nebius-ai-cloud.md), B, 67.2/100. - Schema & documentation: [Replicate Deployments](https://www.anchorterminal.com/tools/replicate-deploy.md), 85/100 on schema & documentation, against 79 for the overall leader. - Agent ergonomics: [Nebius AI Cloud](https://www.anchorterminal.com/tools/nebius-ai-cloud.md), 77/100 on agent ergonomics, against 67 for the overall leader. - Security & auth: [Hugging Face Inference Endpoints](https://www.anchorterminal.com/tools/hugging-face-inference-endpoints.md), 83/100 on security & auth, against 76 for the overall leader. - Transparency & trust: [Replicate Deployments](https://www.anchorterminal.com/tools/replicate-deploy.md), 78/100 on transparency & trust, against 72 for the overall leader. - Lowest paid price per vCPU hour: [Northflank](https://www.anchorterminal.com/tools/northflank.md), $0.0167 per vCPU hour, the lowest of the 7 listings here with a paid price in this unit (free allowances aside). Also [Cerebrium](https://www.anchorterminal.com/tools/cerebrium.md), $0.0236 per vCPU hour. - A hosted MCP endpoint: [Thunder Compute](https://www.anchorterminal.com/tools/thunder-compute.md), remote MCP server, nothing to install. - Self-hosting under an open licence: [Beam](https://www.anchorterminal.com/tools/beam.md), self-hosted, AGPL-3 licence. ## How to choose - Cold start from zero: Check the cold start time from zero and whether it includes loading model weights, because an agent scaled to zero waits through it before answering. - Billing per second or hour: Check whether billing is per second or per hour, and whether idle warm instances are charged, since that sets the real cost of bursty agent traffic. - GPU types and memory: Check which GPU types each region has and how much memory each type holds, because a model that does not fit in memory cannot run. - Autoscaling and scale to zero: Check how quickly the endpoint scales up under load and whether it can scale to zero, because a slow scale-up makes concurrent agent calls queue or time out. - How the benchmark tests this category: The same model deployed as an endpoint on each platform, called cold and warm, then scaled to zero. We time cold starts, check the scaling and add up the cost per GPU-hour. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/gpu-compute/index.md