# GPU and serverless compute for AI workloads > 9 GPU & serverless compute listings ranked by the Anchor benchmark. Leader Modal Sandboxes (BB). Clouds that run your own models and jobs on GPUs by the second, as serverless functions, endpoints or rented machines. Compared on GPU types and price per hour, cold starts, scaling and what you have to package. - Canonical: https://www.anchorterminal.com/categories/gpu-compute - Markdown: https://www.anchorterminal.com/categories/gpu-compute.md (~2,750 tokens) - Slim: https://www.anchorterminal.com/categories/gpu-compute.min.md (~480 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/gpu-compute.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-05 Clouds that run your own models and jobs on GPUs by the second, as serverless functions, endpoints or rented machines. Compared on GPU types and price per hour, cold starts, scaling and what you have to package. - Tools ranked: 9 · agent-ready (BB or better): 1 · accept x402: 0 · hosted endpoints: 7 · desk reviews by the panel: 24 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers - https://letme.dev/compute.gpu picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 33 | Modal Sandboxes | Modal | SDK + MCP | Sandboxes | BB | 75.6 | medium | no | API key | local | 3.3/5 (8) | https://www.anchorterminal.com/tools/modal-sandboxes.md | | 157 | Baseten | Baseten | HTTP API | GPU compute | B | 66.7 | medium | no | API key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/baseten.md | | 195 | Modal | Modal | Model platform | GPU compute | B | 63.8 | medium | no | API key | local | 4/5 (2) | https://www.anchorterminal.com/tools/modal.md | | 197 | Replicate Deployments | Replicate | HTTP API | GPU compute | B | 63.7 | medium | no | API key | hosted + local | 3/5 (2) | https://www.anchorterminal.com/tools/replicate-deploy.md | | 224 | Northflank | Northflank | Model platform | GPU compute | C | 61.8 | medium | no | Token | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/northflank.md | | 313 | Beam | Beam | Model platform | GPU compute | C | 55.5 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/beam.md | | 329 | Runpod | Runpod | HTTP API | GPU compute | D | 53.7 | medium | no | OAuth or key | hosted + local | 3/5 (2) | https://www.anchorterminal.com/tools/runpod.md | | 363 | Lambda Cloud | Lambda | HTTP API | GPU compute | D | 50.1 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/lambda.md | | 389 | Koyeb | Koyeb | Model platform | GPU compute | D | 47 | medium | no | Token | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/koyeb.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 33. Modal Sandboxes, BB (75.6) Modal's sandboxed compute environments for running code, with SDK access, GPU support and filesystem snapshots. GPU sandboxes at the same per-second rates as the rest of Modal. No REST API, and the JavaScript and Go SDKs are beta. - Page: https://www.anchorterminal.com/tools/modal-sandboxes · Markdown: https://www.anchorterminal.com/tools/modal-sandboxes.md · JSON: https://www.anchorterminal.com/api/v1/tools/modal-sandboxes.json - Capabilities: sandbox.code, sandbox.fs, sandbox.persist, sandbox.gpu ### 157. Baseten, B (66.7) Dedicated model deployments packaged with the open-source Truss framework and served behind a per-model HTTPS endpoint, with autoscaling from zero replicas, async inference, a management API and per-minute GPU billing from T4 to B200. Team API keys scoped to inference-only, metrics-only or a single environment or model, plus a Viewer role since 1 September 2026. H100 at $6.50 and A100 at $4.00 an hour, and start-up and idle replica time are billed. - Page: https://www.anchorterminal.com/tools/baseten · Markdown: https://www.anchorterminal.com/tools/baseten.md · JSON: https://www.anchorterminal.com/api/v1/tools/baseten.json - Capabilities: compute.gpu, compute.endpoints, compute.serverless, compute.containers · endpoint: `https://api.baseten.co` ### 195. Modal, B (63.8) Serverless functions, web endpoints, servers and GPU jobs from a Python decorator, with JavaScript and Go SDKs. Scale to zero by default, per-second billing and about one-second container boots. No REST API or OpenAPI spec for deploying or invoking Functions. - Page: https://www.anchorterminal.com/tools/modal · Markdown: https://www.anchorterminal.com/tools/modal.md · JSON: https://www.anchorterminal.com/api/v1/tools/modal.json - Capabilities: compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers ### 197. Replicate Deployments, B (63.7) Replicate's service for deploying and running custom models. OpenAPI file, llms.txt and an MCP server with a two-tool code mode. Private instances bill set-up and idle time, H100 at $5.49 an hour. - Page: https://www.anchorterminal.com/tools/replicate-deploy · Markdown: https://www.anchorterminal.com/tools/replicate-deploy.md · JSON: https://www.anchorterminal.com/api/v1/tools/replicate-deploy.json - Capabilities: compute.gpu, compute.endpoints, compute.serverless, compute.containers · endpoint: `https://api.replicate.com/v1` ### 224. Northflank, C (61.8) Platform for deploying services, jobs and databases from Git or container images, in managed or customer-owned infrastructure. Supports GPU workloads and sandboxes. OpenAPI 3.0 with over 100 paths, enums and `per_page`, `page` and `cursor` on every list. No scale to zero for services; minimum instances must be at least 1. - Page: https://www.anchorterminal.com/tools/northflank · Markdown: https://www.anchorterminal.com/tools/northflank.md · JSON: https://www.anchorterminal.com/api/v1/tools/northflank.json - Capabilities: compute.gpu, compute.containers, compute.batch, compute.endpoints · endpoint: `https://api.northflank.com/v1` ### 313. Beam, C (55.5) Serverless GPU endpoints, task queues, functions, pods and sandboxes from Python decorators, on the open-source beta9 runtime. Per-millisecond billing with cold starts and image pulls free, H100 PCIe at $3.50 and RTX 4090 at $0.69 an hour. No published request rate limits, 429 handling or SLA. - Page: https://www.anchorterminal.com/tools/beam · Markdown: https://www.anchorterminal.com/tools/beam.md · JSON: https://www.anchorterminal.com/api/v1/tools/beam.json - Capabilities: compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers · endpoint: `https://app.beam.cloud/api/v1` ### 329. Runpod, D (53.7) Serverless GPU endpoints, queue-based or load-balanced, and rented GPU Pods, billed per second from prepaid credit. Per-second billing across more than a dozen serverless GPU classes, H100 at $4.79 and A100 80 GB at $2.72 an hour. Data-centre outages of 6 to 24 hours in each of July, August and September 2026. - Page: https://www.anchorterminal.com/tools/runpod · Markdown: https://www.anchorterminal.com/tools/runpod.md · JSON: https://www.anchorterminal.com/api/v1/tools/runpod.json - Capabilities: compute.gpu, compute.serverless, compute.endpoints, compute.batch, compute.containers · endpoint: `https://api.runpod.ai/v2` ### 363. Lambda Cloud, D (50.1) On-demand GPU virtual machines and clusters, with an API for provisioning compute and persistent storage. H100 SXM at $3.99 and B200 at $6.69 an hour with per-minute billing. No scale to zero, autoscaling or endpoints; an idle VM bills until terminated. - Page: https://www.anchorterminal.com/tools/lambda · Markdown: https://www.anchorterminal.com/tools/lambda.md · JSON: https://www.anchorterminal.com/api/v1/tools/lambda.json - Capabilities: compute.gpu, compute.containers · endpoint: `https://cloud.lambda.ai/api/v1` ### 389. Koyeb, D (47) Serverless platform for deploying applications and containers on CPU or GPU instances, with autoscaling and an API. H100 at $2.50 and H200 at $3.00 an hour, billed per second. Changelog silent since 27 February 2026 and no CLI release since 12 May 2026. - Page: https://www.anchorterminal.com/tools/koyeb · Markdown: https://www.anchorterminal.com/tools/koyeb.md · JSON: https://www.anchorterminal.com/api/v1/tools/koyeb.json - Capabilities: compute.gpu, compute.serverless, compute.endpoints, compute.containers · endpoint: `https://app.koyeb.com/v1` ## How we test this category The same model deployed as an endpoint on each platform, called cold and warm, then scaled to zero. We time cold starts, check the scaling and add up the cost per GPU-hour. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence. ## Indexed, not reviewed (6) Sorted into this category from public catalogues, with facts and our own checks but no score, grade or rank (https://www.anchorterminal.com/indexed/index.md). | Listing | Kind | What it does | Why it's here | | --- | --- | --- | --- | | [3DOptix](https://www.anchorterminal.com/tools/3doptix-optical-design.md) | MCP server | Optical design, simulation and analysis with GPU-powered ray tracing. Import from Zemax and CAD. | vendor's own | | [AmpleRun GPU rentals](https://www.anchorterminal.com/tools/amplerun-gpu-rentals.md) | MCP server | Find, price, rent and stop GPUs on AmpleRun. Paid in USDC or USDT on Base. Flat 5% fee. | vendor's own | | [FitLLM](https://www.anchorterminal.com/tools/fitllm.md) | MCP server | Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only. | vendor's own | | [framebench](https://www.anchorterminal.com/tools/framebench.md) | MCP server | Estimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks. | vendor's own | | [prismnetwork.tech MCP server](https://www.anchorterminal.com/tools/prismnetwork-mcp.md) | MCP server | Rent real NVIDIA GPUs from your agent. Browse with no wallet; pay per second onchain. | vendor's own | | [Zhijiangyun Cloud (智匠云)](https://www.anchorterminal.com/tools/artibot-zhijiangyun.md) | MCP server | Robot embodied-AI cloud for GPU training, inference, benchmarks, and robot data collection. | vendor's own |