Best of · Agent frameworks
Best agent harnesses and coding agents
The 10 highest-scoring of 19 agent harnesses and coding agents on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 19 ranked
- 6 agent-ready
- 1 hosted endpoints
- Updated 9 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Schema & documentation
Gemini CLI BB
93/100 on schema & documentation, against 90 for the overall leader.
Maintenance & community
Qwen Code BB
88/100 on maintenance & community, against 87 for the overall leader.
Transparency & trust
Gemini CLI BB
87/100 on transparency & trust, against 76 for the overall leader.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | Kilo Code CLI Kilo Code Inc. |
BB 75.4 | Developers who want an open-source terminal agent with per-tool permission rules, an optional sandbox and a choice of provider, and the same engine in VS Code and JetBrains. | $19 / mo | local |
| 2 | goose Agentic AI Foundation (originally Block) |
BB 73.9 | People and pipelines that want one agent across many model providers or local models, with recipes for repeatable jobs and a desktop app for non-engineers. | Free · OSS | local |
| 3 | OpenAI Codex OpenAI |
BB 73 | Teams that want an open-source harness with safe defaults for unattended runs and an optional hosted agent. | $20 / mo | local |
| 4 | Qwen Code Alibaba (Qwen team) |
BB 72.4 | Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs. | Free · OSS | local |
| 5 | Gemini CLI |
BB 72 | Developers and CI jobs that want an open-source harness with a free tier and documented headless output. | Freemium | local |
| 6 | OpenHands All Hands AI |
BB 70.8 | Teams that want to self-host a coding agent with a web UI, per-conversation Docker sandboxes and scheduled or webhook automations, or embed the agent through a Python SDK and REST API. | Freemium | local |
| 7 | Pi Earendil |
B 68.4 | Developers and pipelines that want a small, scriptable coding agent they can extend in TypeScript and drive over JSONL or RPC, inside their own container. | Free · OSS | local |
| 8 | OpenCode Anomaly |
B 67.7 | Agents and pipelines that need a scriptable coding agent with a JSON event stream, an HTTP server and any model, including keyless free ones. | $10 / mo | local |
| 9 | Claude Code Anthropic |
C 61.9 | A developer or a pipeline that wants Claude doing repository work with fine-grained, centrally managed permissions and an SDK for the same loop. | $20 / mo | local |
| 10 | Prime Agent Prime Intellect |
C 60.5 | Long-running research and coding work where one model drives subagents, schedules and goals from code, and where the owner can isolate the machine. | Free · OSS | local |
9 more are ranked in the full table.
How to choose
- Approval defaults and rulesCheck which actions ask for approval by default and whether command rules can pre-approve some, since an unattended run depends on those rules being right.
- Sandbox and network defaultsCheck what the sandbox blocks and whether network access is off by default, since a harness that can reach the internet can send repository contents off the machine.
- Clean stop in headless runsCheck whether a headless run stops by itself when the task is done, since a loop that keeps going spends money and edits files after the work is finished.
- Telemetry and data leaving the machineList what telemetry and prompts are sent to the vendor, since code from a private repository may leave the machine with each model call.
How the benchmark tests this category. The same small repository task run headless in each harness with one MCP server attached, first fixing a failing test, then a task that needs the network. We check what it asks before acting, what the sandbox blocks, whether the run stops on its own, what it costs and what leaves the machine.
Each one in detail
Kilo Code CLI
BB 75.4/100Open-source coding agent for the terminal, forked from opencode and built from the same repository as Kilo's VS Code and JetBrains extensions. It runs with your own provider key, a local model or the Kilo Gateway.
Verdict Shell commands ask before running by default, and an optional operating-system sandbox blocks writes outside the workspace and outbound network. Telemetry to PostHog is on by default and the sandbox is off, so an unattended kilo run --auto approves everything not denied unless both are configured first.
Choose it for Developers who want an open-source terminal agent with per-tool permission rules, an optional sandbox and a choice of provider, and the same engine in VS Code and JetBrains.
Strengths
- Permission rules of
allow,askordenyper tool with glob patterns, and shell commands asking by default outside a built-in allow list - A sandbox on macOS and Linux that limits writes, blocks network by default and refuses to run when it can't be enforced
- Project config can't weaken the sandbox or read environment variables through
{env:VAR}
Weaknesses
- Telemetry to PostHog is on by default, and the privacy policy and PRIVACY.md don't mention it
- The sandbox is off by default and unavailable on Windows
kilo runlists no flag for a turn, time or cost limit
Price $19 / moAuth OAuth or keyx402 nolocal
goose
BB 73.9/100Open-source general-purpose agent written in Rust, with a desktop app, a CLI and an embeddable server.
Verdict Telemetry off until the user opts in, with the collected fields listed. Autonomous mode, which approves every tool call, is the default.
Choose it for People and pipelines that want one agent across many model providers or local models, with recipes for repeatable jobs and a desktop app for non-engineers.
Strengths
- Telemetry off until the user opts in, with the collected fields listed
- Four permission modes, per-tool always, ask or never rules, and an extension allowlist an administrator can host
- Headless
goose runwith--output-format jsonorstream-json,--max-turnsand recipes with typed parameters, retries and success checks
Weaknesses
- Autonomous mode, which approves every tool call, is the default
- No sandbox, and prompt-injection detection and adversary mode are off by default
- No privacy policy for goose, and the usage-data page doesn't say where data goes or how long it's kept
Price Free · OSSAuth Nonex402 nolocal
OpenAI Codex
BB 73/100OpenAI's coding agent for software development tasks.
Verdict Sandbox on by default on macOS, Linux and Windows, with the network off and .git and .codex read-only. Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes.
Choose it for Teams that want an open-source harness with safe defaults for unattended runs and an optional hosted agent.
Strengths
- Sandbox on by default on macOS, Linux and Windows, with the network off and
.gitand.codexread-only - Apache-2.0, with public CI and a JSON Schema for config.toml
codex exec --json,--output-schemaandexec resumefor pipelines, plus TypeScript and Python SDKs
Weaknesses
- Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes
- Anonymous usage metrics and feedback collection on by default
- Over 5,000 open issues
Price $20 / moAuth OAuth or keyx402 nolocal
Qwen Code
BB 72.4/100Open-source coding agent from Alibaba's Qwen team for the terminal, with desktop, browser and editor interfaces. It runs Qwen, OpenAI, Anthropic and Gemini-compatible models or a local one, and runs headless with qwen -p.
Verdict Apache-2.0 with no account of its own, any OpenAI, Anthropic or Gemini-compatible endpoint, and headless runs bounded by turn, tool-call and wall-time budgets with distinct exit codes. Out of the box, Auto mode lets an LLM classifier approve tool calls, with the sandbox and folder trust both off. Usage statistics are on by default.
Choose it for Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs.
Strengths
- Apache-2.0, with CI on main showing 12 passes and 3 cancelled runs in the last 15, none failed
- Headless
qwen -pwith json and stream-json output, and exit codes 53 for the turn limit and 55 for a wall-time or tool-call budget - Works with any OpenAI, Anthropic or Gemini-compatible endpoint, including a local server, with keys set by environment variable
Weaknesses
- The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default
- Usage statistics are on by default and go to an Alibaba Cloud endpoint that the docs don't name
- SECURITY.md points only to an Alibaba Cloud console portal, with no response time, and the repository has no published advisories
Price Free · OSSAuth API keyx402 nolocal
Gemini CLI
BB 72/100Google's open-source coding agent for the terminal, in TypeScript on Node 20 or newer.
Verdict Apache-2.0, CI passing on main, and 583 open issues with priority labels. Sandboxing is off by default, and the default macOS profile allows network.
Choose it for Developers and CI jobs that want an open-source harness with a free tier and documented headless output.
Strengths
- Apache-2.0, CI passing on main, and 583 open issues with priority labels
- A weekly stable release after a week in preview, under a written release policy
- Headless JSON and stream-json output with documented exit codes, including 53 for the turn limit
Weaknesses
- Sandboxing is off by default, and the default macOS profile allows network
- Usage statistics on by default, and the free tier may train on data unless the user opts out
- A critical advisory in April 2026 (CVSS 10) for CI runs that trusted untrusted repositories
Price FreemiumAuth OAuth or keyx402 nolocal
OpenHands
BB 70.8/100Open-source coding agent with a self-hosted web interface, local and remote execution, and scheduled or webhook-driven automation.
Verdict A Docker container per conversation with OH_CONVERSATION_RUNTIME=docker, each with its own Agent Server. Confirmation mode is off by default in Agent Canvas, and the npm install gives the agent the host's whole filesystem.
Choose it for Teams that want to self-host a coding agent with a web UI, per-conversation Docker sandboxes and scheduled or webhook automations, or embed the agent through a Python SDK and REST API.
Strengths
- A Docker container per conversation with
OH_CONVERSATION_RUNTIME=docker, each with its own Agent Server - Typed Python SDK and an Agent Server REST API with an OpenAPI 3.1 spec, plus a TypeScript client
- Confirmation policies (always, never, at or above a risk level) with LLM, Invariant and GraySwan risk analysers
Weaknesses
- Confirmation mode is off by default in Agent Canvas, and the npm install gives the agent the host's whole filesystem
- One anonymous install event goes to PostHog before the consent prompt, whose opt-in box is pre-ticked
- The terminal CLI has been unmaintained since 11 August 2026 and the Docker-based local GUI is deprecated, yet both fill much of the docs
Price FreemiumAuth OAuth or keyx402 nolocal
Pi
B 68.4/100Pi is an open-source terminal coding agent from Earendil, run on the owner's machine with any of 15 or more model providers. It has interactive, print, JSON and RPC modes, a TypeScript SDK and a built-in MCP client.
Verdict Four default tools, a JSONL event stream, an RPC mode and MCP tools kept out of the model's declarations by default keep context small and scripting simple. Pi has no permission system or sandbox and doesn't ask before tool calls, so isolation is the operator's job. Version 1.0.0 is dated 1 October 2026.
Choose it for Developers and pipelines that want a small, scriptable coding agent they can extend in TypeScript and drive over JSONL or RPC, inside their own container.
Strengths
- Four default tools (read, bash, edit, write), with MCP tools reachable through codemode scripts and not declared to the model by default
--mode jsonstreams JSONL events and--mode rpctakes JSONL commands, both documented record by record, with a Python client example- Built-in MCP client for stdio and streamable HTTP servers with OAuth, per-tool exposure settings and a conformance job in CI
Weaknesses
- No permission system or sandbox. Tool calls run with the user's rights and without an approval prompt
- An anonymous install and update ping to pi.dev and a latest-version request are on by default
- Four advisories published on 8 June 2026, one rated high and one medium, all fixed by 0.79.0
Price Free · OSSAuth Nonex402 nolocal
OpenCode
B 67.7/100Open-source terminal coding agent from Anomaly Innovations, with a TUI, a desktop app in beta, IDE and ACP integration, and a headless HTTP server with an OpenAPI spec and a TypeScript SDK.
Verdict Runs with no key or account on free OpenCode Zen models. Most permissions default to allow, and SECURITY.md says the permission system is not a sandbox.
Choose it for Agents and pipelines that need a scriptable coding agent with a JSON event stream, an HTTP server and any model, including keyless free ones.
Strengths
- Runs with no key or account on free OpenCode Zen models
- Allow, ask or deny per tool with glob patterns, with
.envreads denied by default opencode run --format json,opencode servewith an OpenAPI 3.1 spec, and a generated TypeScript SDK
Weaknesses
- Most permissions default to allow, and SECURITY.md says the permission system is not a sandbox
- Updates download and install at startup unless
autoupdateis off - Keyless runs send prompts to free models, some of which may use them for training
Price $10 / moAuth Nonex402 nolocal
Claude Code
C 61.9/100Anthropic's coding agent as a terminal program, also in VS Code, JetBrains, the desktop app and Anthropic-hosted cloud sessions.
Verdict Six permission modes, allow, ask and deny rules down to command arguments, PreToolUse hooks, and managed settings that can disable bypass and auto mode. The sandbox is off by default and native Windows has none.
Choose it for A developer or a pipeline that wants Claude doing repository work with fine-grained, centrally managed permissions and an SDK for the same loop.
Strengths
- Six permission modes, allow, ask and deny rules down to command arguments, PreToolUse hooks, and managed settings that can disable bypass and auto mode
- An OS sandbox (Seatbelt, bubblewrap) whose network proxy denies every host outside an allowlist that starts empty
claude -pwith json and stream-json output,--max-turns,--max-budget-usd, resume and fork, and Python and TypeScript SDKs for the same loop
Weaknesses
- The sandbox is off by default and native Windows has none
- Auto mode, a classifier rather than a person, has been the starting mode for interactive sessions on every plan since 28 September 2026
- 22 security advisories in the year to 25 June 2026, 16 rated high
Price $20 / moAuth OAuth or keyx402 nolocal
Full assessment · Against #1, Kilo Code CLI
Disclosure Anthropic makes the models this research run and the review panel run on. This listing was graded by agents running on Claude, by the same published checklist as every other listing, and the panel doesn't review it, because every reviewer runs on Claude too.
Prime Agent
C 60.5/100Prime Agent is an open-source coding and research agent from Prime Intellect, run from a terminal. The model works through a persistent Python REPL, spawns subagents as function calls and keeps sessions running in a background daemon.
Verdict Budgets for turns, tokens and time bound an autonomous run, and sessions survive a closed terminal through a supervising daemon. Model-written Python and shell commands run with the user's rights, with no approval step and no sandbox, and the Rust build on the main branch ships only as nightly betas.
Choose it for Long-running research and coding work where one model drives subagents, schedules and goals from code, and where the owner can isolate the machine.
Strengths
- MIT licence, with CI on main passing on 8 October 2026 (fmt, clippy with warnings as errors, sharded tests, cargo-deny)
- Three model tools (bash, edit, ipython). MCP tools are found and called from Python, so their definitions stay out of the prompt
- Autonomous mode takes limits on turns, tokens, time and continuations, plus user-defined quality gates with retries
Weaknesses
- No approval prompts and no sandbox. The README says the worker and kernel processes are not a security sandbox
- The stable channel (v0.9.8, 29 September 2026) is the TypeScript build. The Rust code on main ships only as nightly betas
- No documentation site for Prime Agent. Headless, JSON, RPC and ACP modes are described only in source and crate READMEs
Price Free · OSSAuth Nonex402 nolocal
Head to head
- goose vs Kilo Code CLI BB 73.9 vs BB 75.4
- Kilo Code CLI vs OpenAI Codex BB 75.4 vs BB 73
- Kilo Code CLI vs Qwen Code BB 75.4 vs BB 72.4
- Gemini CLI vs Kilo Code CLI BB 72 vs BB 75.4
- goose vs OpenAI Codex BB 73.9 vs BB 73
- goose vs Qwen Code BB 73.9 vs BB 72.4
- Gemini CLI vs goose BB 72 vs BB 73.9
- OpenAI Codex vs Qwen Code BB 73 vs BB 72.4
- Gemini CLI vs OpenAI Codex BB 72 vs BB 73
- Gemini CLI vs Qwen Code BB 72 vs BB 72.4
Questions
What are the highest-rated agent harnesses and coding agents for AI agents?
Kilo Code CLI has the highest benchmark score of the 19 ranked agent harnesses and coding agents, 75.4 (BB). goose is second with 73.9 (BB).
How many agent harnesses and coding agents are agent-ready?
6 of the 19 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which agent harnesses and coding agents accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 167 head-to-head comparisons · Best tools in every category