Category · Models & inference
GPU and serverless compute for AI workloads
Clouds that run your own models and jobs on GPUs by the second, as serverless functions, endpoints or rented machines. Compared on GPU types and price per hour, cold starts, scaling and what you have to package.
Capability keys compute.gpu · compute.serverless · compute.endpoints · compute.batch · compute.containers · All tools
letme.dev/compute.gpu picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 33 | Modal SandboxesModal · SDK + MCP | Sandboxes | BB | 75.6 | 3.3 (8) | n/a | n/a | $0.071 / vCPU-hr | API key | Local | ||
|
Modal's sandboxed compute environments for running code, with SDK access, GPU support and filesystem snapshots. Top strength GPU sandboxes at the same per-second rates as the rest of Modal Top weakness No REST API, and the JavaScript and Go SDKs are beta |
||||||||||||
| 157 | BasetenBaseten · HTTP API | GPU compute | B | 66.7 | 3.5 (2) | n/a | n/a | Pay per use | API key | Hosted | ||
|
Dedicated model deployments packaged with the open-source Truss framework and served behind a per-model HTTPS endpoint, with autoscaling from zero replicas, async inference, a management API and per-minute GPU billing from T4 to B200. Top strength Team API keys scoped to inference-only, metrics-only or a single environment or model, plus a Viewer role since 1 September 2026 Top weakness H100 at $6.50 and A100 at $4.00 an hour, and start-up and idle replica time are billed |
||||||||||||
| 195 | ModalModal · Model platform | GPU compute | B | 63.8 | 4.0 (2) | n/a | n/a | $250 / mo | API key | Local | ||
|
Serverless functions, web endpoints, servers and GPU jobs from a Python decorator, with JavaScript and Go SDKs. Top strength Scale to zero by default, per-second billing and about one-second container boots Top weakness No REST API or OpenAPI spec for deploying or invoking Functions |
||||||||||||
| 197 | Replicate DeploymentsReplicate · HTTP API | GPU compute | B | 63.7 | 3.0 (2) | n/a | n/a | Pay per use | API key | Hosted + local | ||
|
Replicate's service for deploying and running custom models. Top strength OpenAPI file, llms.txt and an MCP server with a two-tool code mode Top weakness Private instances bill set-up and idle time, H100 at $5.49 an hour |
||||||||||||
| 224 | NorthflankNorthflank · Model platform | GPU compute | C | 61.8 | 3.5 (2) | n/a | n/a | $0.0167 / vCPU-hr | Token | Hosted | ||
|
Platform for deploying services, jobs and databases from Git or container images, in managed or customer-owned infrastructure. Supports GPU workloads and sandboxes. Top strength OpenAPI 3.0 with over 100 paths, enums and Top weakness No scale to zero for services; minimum instances must be at least 1 |
||||||||||||
| 313 | BeamBeam · Model platform | GPU compute | C | 55.5 | 3.0 (2) | n/a | n/a | $0.045 / vCPU-hr | API key | Hosted | ||
|
Serverless GPU endpoints, task queues, functions, pods and sandboxes from Python decorators, on the open-source beta9 runtime. Top strength Per-millisecond billing with cold starts and image pulls free, H100 PCIe at $3.50 and RTX 4090 at $0.69 an hour Top weakness No published request rate limits, 429 handling or SLA |
||||||||||||
| 329 | RunpodRunpod · HTTP API | GPU compute | D | 53.7 | 3.0 (2) | n/a | n/a | Pay per use | OAuth or key | Hosted + local | ||
|
Serverless GPU endpoints, queue-based or load-balanced, and rented GPU Pods, billed per second from prepaid credit. Top strength Per-second billing across more than a dozen serverless GPU classes, H100 at $4.79 and A100 80 GB at $2.72 an hour Top weakness Data-centre outages of 6 to 24 hours in each of July, August and September 2026 |
||||||||||||
| 363 | Lambda CloudLambda · HTTP API | GPU compute | D | 50.1 | 3.0 (2) | n/a | n/a | Pay per use | API key | Hosted | ||
|
On-demand GPU virtual machines and clusters, with an API for provisioning compute and persistent storage. Top strength H100 SXM at $3.99 and B200 at $6.69 an hour with per-minute billing Top weakness No scale to zero, autoscaling or endpoints; an idle VM bills until terminated |
||||||||||||
| 389 | KoyebKoyeb · Model platform | GPU compute | D | 47 | 3.5 (2) | n/a | n/a | $29 / mo | Token | Hosted | ||
|
Serverless platform for deploying applications and containers on CPU or GPU instances, with autoscaling and an API. Top strength H100 at $2.50 and H200 at $3.00 an hour, billed per second Top weakness Changelog silent since 27 February 2026 and no CLI release since 12 May 2026 |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
Indexed, not reviewed (6)
Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.
| Listing | Kind | What it does | Why it's here |
|---|---|---|---|
| 3DOptix 3doptix.com | MCP server | Optical design, simulation and analysis with GPU-powered ray tracing. Import from Zemax and CAD. | vendor's own |
| AmpleRun GPU rentals amplerun.com | MCP server | Find, price, rent and stop GPUs on AmpleRun. Paid in USDC or USDT on Base. Flat 5% fee. | vendor's own |
| FitLLM fitllm.run | MCP server | Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only. | vendor's own |
| framebench framebench.app | MCP server | Estimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks. | vendor's own |
| prismnetwork.tech MCP server prismnetwork.tech | MCP server | Rent real NVIDIA GPUs from your agent. Browse with no wallet; pay per second onchain. | vendor's own |
| Zhijiangyun Cloud (智匠云) artibot.cloud | MCP server | Robot embodied-AI cloud for GPU training, inference, benchmarks, and robot data collection. | vendor's own |
How we test this category
The same model deployed as an endpoint on each platform, called cold and warm, then scaled to zero. We time cold starts, check the scaling and add up the cost per GPU-hour. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.







