Category · Models & inference

GPU and serverless compute for AI workloads

Clouds that run your own models and jobs on GPUs by the second, as serverless functions, endpoints or rented machines. Compared on GPU types and price per hour, cold starts, scaling and what you have to package.

Capability keys compute.gpu · compute.serverless · compute.endpoints · compute.batch · compute.containers · All tools

letme.dev/compute.gpu picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

9listings graded
1agent-ready (BB+)
24desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Mode
Auth
Pricing
Status
9 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
33 Modal SandboxesModal · SDK + MCP Sandboxes BB 75.6 3.3 (8) $0.071 / vCPU-hr
157 BasetenBaseten · HTTP API GPU compute B 66.7 3.5 (2) Pay per use
195 ModalModal · Model platform GPU compute B 63.8 4.0 (2) $250 / mo
197 Replicate DeploymentsReplicate · HTTP API GPU compute B 63.7 3.0 (2) Pay per use
224 NorthflankNorthflank · Model platform GPU compute C 61.8 3.5 (2) $0.0167 / vCPU-hr
313 BeamBeam · Model platform GPU compute C 55.5 3.0 (2) $0.045 / vCPU-hr
329 RunpodRunpod · HTTP API GPU compute D 53.7 3.0 (2) Pay per use
363 Lambda CloudLambda · HTTP API GPU compute D 50.1 3.0 (2) Pay per use
389 KoyebKoyeb · Model platform GPU compute D 47 3.5 (2) $29 / mo

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

Indexed, not reviewed (6)

Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.

ListingKindWhat it doesWhy it's here
3DOptix
3doptix.com
MCP serverOptical design, simulation and analysis with GPU-powered ray tracing. Import from Zemax and CAD.vendor's own
AmpleRun GPU rentals
amplerun.com
MCP serverFind, price, rent and stop GPUs on AmpleRun. Paid in USDC or USDT on Base. Flat 5% fee.vendor's own
FitLLM
fitllm.run
MCP serverWill this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.vendor's own
framebench
framebench.app
MCP serverEstimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks.vendor's own
prismnetwork.tech MCP server
prismnetwork.tech
MCP serverRent real NVIDIA GPUs from your agent. Browse with no wallet; pay per second onchain.vendor's own
Zhijiangyun Cloud (智匠云)
artibot.cloud
MCP serverRobot embodied-AI cloud for GPU training, inference, benchmarks, and robot data collection.vendor's own

How we test this category

The same model deployed as an endpoint on each platform, called cold and warm, then scaled to zero. We time cold starts, check the scaling and add up the cost per GPU-hour. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.