Category · Agent frameworks

Agent harnesses and coding agents

Finished programs that run the agent loop for a person or a pipeline, in a terminal, an editor or the vendor's cloud. They plan, call tools, edit files and run commands, where a framework is a library you build that loop with. Compared on what they ask before acting, what the sandbox and the network allow by default, MCP support, headless use and the telemetry they send.

Capability keys agent.harness · agent.mcp-client · agent.multi-agent · All tools

letme.dev/agent.harness picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

10listings graded
4agent-ready (BB+)
18desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Auth
Pricing
Status
10 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
52 gooseAgentic AI Foundation (originally Block) · Agent harness Harnesses BB 73.9 2.5 (2) Free · OSS
58 OpenAI CodexOpenAI · Agent harness Harnesses BB 73.4 3.0 (2) $20 / mo
72 Gemini CLIGoogle · Agent harness Harnesses BB 72.3 3.5 (2) Freemium
92 OpenHandsAll Hands AI · Agent harness Harnesses BB 70.9 2.5 (2) Freemium
134 OpenCodeAnomaly · Agent harness Harnesses B 68 2.0 (2) $10 / mo
222 Claude CodeAnthropic · Agent harness Harnesses B 62.2 none $20 / mo
239 ClineCline Bot Inc. · Agent harness Harnesses C 60.8 2.0 (2) $9.99 / mo
286 GitHub Copilot CLIGitHub · Agent harness Harnesses C 57.9 2.5 (2) $0.01 / credit
385 AiderAider AI LLC · Agent harness Harnesses D 47.1 1.0 (2) Free · OSS
441 Cursor CLICursor · Agent harness Harnesses F 35.8 1.5 (2) $20 / mo

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

How we test this category

The same small repository task run headless in each harness with one MCP server attached, first fixing a failing test, then a task that needs the network. We check what it asks before acting, what the sandbox blocks, whether the run stops on its own, what it costs and what leaves the machine. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.