# Best agent harnesses and coding agents > Kilo Code CLI (BB), goose (BB) and OpenAI Codex (BB) lead the 19 ranked agent harnesses and coding agents. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Canonical: https://www.anchorterminal.com/best/agent-harnesses/ - Markdown: https://www.anchorterminal.com/best/agent-harnesses/index.md (~5,900 tokens) - Slim: https://www.anchorterminal.com/best/agent-harnesses/index.min.md (~1,480 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/best/agent-harnesses/index.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 The 10 highest-scoring of 19 agent harnesses and coding agents on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 19 · agent-ready (BB or better): 6 · accept x402: 0 · hosted endpoints: 1 - Full ranked table: https://www.anchorterminal.com/categories/agent-harnesses.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/agent-harnesses/index.md (167) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [Kilo Code CLI](https://www.anchorterminal.com/tools/kilo-code-cli.md) | BB | 75.4 | Developers who want an open-source terminal agent with per-tool permission rules, an optional sandbox and a choice of provider, and the same engine in VS Code and JetBrains. | $19 / mo | local | | 2 | [goose](https://www.anchorterminal.com/tools/goose.md) | BB | 73.9 | People and pipelines that want one agent across many model providers or local models, with recipes for repeatable jobs and a desktop app for non-engineers. | Free · OSS | local | | 3 | [OpenAI Codex](https://www.anchorterminal.com/tools/openai-codex.md) | BB | 73 | Teams that want an open-source harness with safe defaults for unattended runs and an optional hosted agent. | $20 / mo | local | | 4 | [Qwen Code](https://www.anchorterminal.com/tools/qwen-code.md) | BB | 72.4 | Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs. | Free · OSS | local | | 5 | [Gemini CLI](https://www.anchorterminal.com/tools/gemini-cli.md) | BB | 72 | Developers and CI jobs that want an open-source harness with a free tier and documented headless output. | Freemium | local | | 6 | [OpenHands](https://www.anchorterminal.com/tools/openhands.md) | BB | 70.8 | Teams that want to self-host a coding agent with a web UI, per-conversation Docker sandboxes and scheduled or webhook automations, or embed the agent through a Python SDK and REST API. | Freemium | local | | 7 | [Pi](https://www.anchorterminal.com/tools/earendil-pi.md) | B | 68.4 | Developers and pipelines that want a small, scriptable coding agent they can extend in TypeScript and drive over JSONL or RPC, inside their own container. | Free · OSS | local | | 8 | [OpenCode](https://www.anchorterminal.com/tools/opencode.md) | B | 67.7 | Agents and pipelines that need a scriptable coding agent with a JSON event stream, an HTTP server and any model, including keyless free ones. | $10 / mo | local | | 9 | [Claude Code](https://www.anchorterminal.com/tools/claude-code.md) | C | 61.9 | A developer or a pipeline that wants Claude doing repository work with fine-grained, centrally managed permissions and an SDK for the same loop. | $20 / mo | local | | 10 | [Prime Agent](https://www.anchorterminal.com/tools/prime-agent.md) | C | 60.5 | Long-running research and coding work where one model drives subagents, schedules and goals from code, and where the owner can isolate the machine. | Free · OSS | local | ## Picks by need - Highest score overall: [Kilo Code CLI](https://www.anchorterminal.com/tools/kilo-code-cli.md), BB, 75.4/100 on the benchmark. Also [goose](https://www.anchorterminal.com/tools/goose.md), BB, 73.9/100. - Schema & documentation: [Gemini CLI](https://www.anchorterminal.com/tools/gemini-cli.md), 93/100 on schema & documentation, against 90 for the overall leader. - Agent ergonomics: [Claude Code](https://www.anchorterminal.com/tools/claude-code.md), 87/100 on agent ergonomics, against 72 for the overall leader. - Security & auth: [OpenAI Codex](https://www.anchorterminal.com/tools/openai-codex.md), 82/100 on security & auth, against 66 for the overall leader. - Maintenance & community: [Qwen Code](https://www.anchorterminal.com/tools/qwen-code.md), 88/100 on maintenance & community, against 87 for the overall leader. - Transparency & trust: [Gemini CLI](https://www.anchorterminal.com/tools/gemini-cli.md), 87/100 on transparency & trust, against 76 for the overall leader. - Self-hosting under an open licence: [OpenHands](https://www.anchorterminal.com/tools/openhands.md), self-hosted, MIT licence. Also [Paperclip](https://www.anchorterminal.com/tools/paperclip.md), self-hosted, MIT licence. ## How to choose - Approval defaults and rules: Check which actions ask for approval by default and whether command rules can pre-approve some, since an unattended run depends on those rules being right. - Sandbox and network defaults: Check what the sandbox blocks and whether network access is off by default, since a harness that can reach the internet can send repository contents off the machine. - Clean stop in headless runs: Check whether a headless run stops by itself when the task is done, since a loop that keeps going spends money and edits files after the work is finished. - Telemetry and data leaving the machine: List what telemetry and prompts are sent to the vendor, since code from a private repository may leave the machine with each model call. - How the benchmark tests this category: The same small repository task run headless in each harness with one MCP server attached, first fixing a failing test, then a task that needs the network. We check what it asks before acting, what the sandbox blocks, whether the run stops on its own, what it costs and what leaves the machine. ## Each one in detail ### 1. Kilo Code CLI, BB 75.4/100 Open-source coding agent for the terminal, forked from opencode and built from the same repository as Kilo's VS Code and JetBrains extensions. It runs with your own provider key, a local model or the Kilo Gateway. - Verdict: Shell commands ask before running by default, and an optional operating-system sandbox blocks writes outside the workspace and outbound network. Telemetry to PostHog is on by default and the sandbox is off, so an unattended `kilo run --auto` approves everything not denied unless both are configured first. - Choose it for: Developers who want an open-source terminal agent with per-tool permission rules, an optional sandbox and a choice of provider, and the same engine in VS Code and JetBrains. - Strength: Permission rules of `allow`, `ask` or `deny` per tool with glob patterns, and shell commands asking by default outside a built-in allow list - Strength: A sandbox on macOS and Linux that limits writes, blocks network by default and refuses to run when it can't be enforced - Strength: Project config can't weaken the sandbox or read environment variables through `{env:VAR}` - Weakness: Telemetry to PostHog is on by default, and the privacy policy and PRIVACY.md don't mention it - Weakness: The sandbox is off by default and unavailable on Windows - Weakness: `kilo run` lists no flag for a turn, time or cost limit - Price: $19 / mo · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/kilo-code-cli.md ### 2. goose, BB 73.9/100 Open-source general-purpose agent written in Rust, with a desktop app, a CLI and an embeddable server. - Verdict: Telemetry off until the user opts in, with the collected fields listed. Autonomous mode, which approves every tool call, is the default. - Choose it for: People and pipelines that want one agent across many model providers or local models, with recipes for repeatable jobs and a desktop app for non-engineers. - Strength: Telemetry off until the user opts in, with the collected fields listed - Strength: Four permission modes, per-tool always, ask or never rules, and an extension allowlist an administrator can host - Strength: Headless `goose run` with `--output-format json` or `stream-json`, `--max-turns` and recipes with typed parameters, retries and success checks - Weakness: Autonomous mode, which approves every tool call, is the default - Weakness: No sandbox, and prompt-injection detection and adversary mode are off by default - Weakness: No privacy policy for goose, and the usage-data page doesn't say where data goes or how long it's kept - Price: Free · OSS · Auth: None · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/goose.md - Against #1: https://www.anchorterminal.com/compare/goose-vs-kilo-code-cli.md ### 3. OpenAI Codex, BB 73/100 OpenAI's coding agent for software development tasks. - Verdict: Sandbox on by default on macOS, Linux and Windows, with the network off and `.git` and `.codex` read-only. Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes. - Choose it for: Teams that want an open-source harness with safe defaults for unattended runs and an optional hosted agent. - Strength: Sandbox on by default on macOS, Linux and Windows, with the network off and `.git` and `.codex` read-only - Strength: Apache-2.0, with public CI and a JSON Schema for config.toml - Strength: `codex exec --json`, `--output-schema` and `exec resume` for pipelines, plus TypeScript and Python SDKs - Weakness: Pre-1.0 at 0.160.0, with a minor every few days and no breaking-change section in release notes - Weakness: Anonymous usage metrics and feedback collection on by default - Weakness: Over 5,000 open issues - Price: $20 / mo · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/openai-codex.md - Against #1: https://www.anchorterminal.com/compare/kilo-code-cli-vs-openai-codex.md ### 4. Qwen Code, BB 72.4/100 Open-source coding agent from Alibaba's Qwen team for the terminal, with desktop, browser and editor interfaces. It runs Qwen, OpenAI, Anthropic and Gemini-compatible models or a local one, and runs headless with `qwen -p`. - Verdict: Apache-2.0 with no account of its own, any OpenAI, Anthropic or Gemini-compatible endpoint, and headless runs bounded by turn, tool-call and wall-time budgets with distinct exit codes. Out of the box, Auto mode lets an LLM classifier approve tool calls, with the sandbox and folder trust both off. Usage statistics are on by default. - Choose it for: Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs. - Strength: Apache-2.0, with CI on main showing 12 passes and 3 cancelled runs in the last 15, none failed - Strength: Headless `qwen -p` with json and stream-json output, and exit codes 53 for the turn limit and 55 for a wall-time or tool-call budget - Strength: Works with any OpenAI, Anthropic or Gemini-compatible endpoint, including a local server, with keys set by environment variable - Weakness: The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default - Weakness: Usage statistics are on by default and go to an Alibaba Cloud endpoint that the docs don't name - Weakness: SECURITY.md points only to an Alibaba Cloud console portal, with no response time, and the repository has no published advisories - Price: Free · OSS · Auth: API key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/qwen-code.md - Against #1: https://www.anchorterminal.com/compare/kilo-code-cli-vs-qwen-code.md ### 5. Gemini CLI, BB 72/100 Google's open-source coding agent for the terminal, in TypeScript on Node 20 or newer. - Verdict: Apache-2.0, CI passing on main, and 583 open issues with priority labels. Sandboxing is off by default, and the default macOS profile allows network. - Choose it for: Developers and CI jobs that want an open-source harness with a free tier and documented headless output. - Strength: Apache-2.0, CI passing on main, and 583 open issues with priority labels - Strength: A weekly stable release after a week in preview, under a written release policy - Strength: Headless JSON and stream-json output with documented exit codes, including 53 for the turn limit - Weakness: Sandboxing is off by default, and the default macOS profile allows network - Weakness: Usage statistics on by default, and the free tier may train on data unless the user opts out - Weakness: A critical advisory in April 2026 (CVSS 10) for CI runs that trusted untrusted repositories - Price: Freemium · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/gemini-cli.md - Against #1: https://www.anchorterminal.com/compare/gemini-cli-vs-kilo-code-cli.md ### 6. OpenHands, BB 70.8/100 Open-source coding agent with a self-hosted web interface, local and remote execution, and scheduled or webhook-driven automation. - Verdict: A Docker container per conversation with `OH_CONVERSATION_RUNTIME=docker`, each with its own Agent Server. Confirmation mode is off by default in Agent Canvas, and the npm install gives the agent the host's whole filesystem. - Choose it for: Teams that want to self-host a coding agent with a web UI, per-conversation Docker sandboxes and scheduled or webhook automations, or embed the agent through a Python SDK and REST API. - Strength: A Docker container per conversation with `OH_CONVERSATION_RUNTIME=docker`, each with its own Agent Server - Strength: Typed Python SDK and an Agent Server REST API with an OpenAPI 3.1 spec, plus a TypeScript client - Strength: Confirmation policies (always, never, at or above a risk level) with LLM, Invariant and GraySwan risk analysers - Weakness: Confirmation mode is off by default in Agent Canvas, and the npm install gives the agent the host's whole filesystem - Weakness: One anonymous install event goes to PostHog before the consent prompt, whose opt-in box is pre-ticked - Weakness: The terminal CLI has been unmaintained since 11 August 2026 and the Docker-based local GUI is deprecated, yet both fill much of the docs - Price: Freemium · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/openhands.md - Against #1: https://www.anchorterminal.com/compare/kilo-code-cli-vs-openhands.md ### 7. Pi, B 68.4/100 Pi is an open-source terminal coding agent from Earendil, run on the owner's machine with any of 15 or more model providers. It has interactive, print, JSON and RPC modes, a TypeScript SDK and a built-in MCP client. - Verdict: Four default tools, a JSONL event stream, an RPC mode and MCP tools kept out of the model's declarations by default keep context small and scripting simple. Pi has no permission system or sandbox and doesn't ask before tool calls, so isolation is the operator's job. Version 1.0.0 is dated 1 October 2026. - Choose it for: Developers and pipelines that want a small, scriptable coding agent they can extend in TypeScript and drive over JSONL or RPC, inside their own container. - Strength: Four default tools (read, bash, edit, write), with MCP tools reachable through codemode scripts and not declared to the model by default - Strength: `--mode json` streams JSONL events and `--mode rpc` takes JSONL commands, both documented record by record, with a Python client example - Strength: Built-in MCP client for stdio and streamable HTTP servers with OAuth, per-tool exposure settings and a conformance job in CI - Weakness: No permission system or sandbox. Tool calls run with the user's rights and without an approval prompt - Weakness: An anonymous install and update ping to pi.dev and a latest-version request are on by default - Weakness: Four advisories published on 8 June 2026, one rated high and one medium, all fixed by 0.79.0 - Price: Free · OSS · Auth: None · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/earendil-pi.md - Against #1: https://www.anchorterminal.com/compare/earendil-pi-vs-kilo-code-cli.md ### 8. OpenCode, B 67.7/100 Open-source terminal coding agent from Anomaly Innovations, with a TUI, a desktop app in beta, IDE and ACP integration, and a headless HTTP server with an OpenAPI spec and a TypeScript SDK. - Verdict: Runs with no key or account on free OpenCode Zen models. Most permissions default to allow, and SECURITY.md says the permission system is not a sandbox. - Choose it for: Agents and pipelines that need a scriptable coding agent with a JSON event stream, an HTTP server and any model, including keyless free ones. - Strength: Runs with no key or account on free OpenCode Zen models - Strength: Allow, ask or deny per tool with glob patterns, with `.env` reads denied by default - Strength: `opencode run --format json`, `opencode serve` with an OpenAPI 3.1 spec, and a generated TypeScript SDK - Weakness: Most permissions default to allow, and SECURITY.md says the permission system is not a sandbox - Weakness: Updates download and install at startup unless `autoupdate` is off - Weakness: Keyless runs send prompts to free models, some of which may use them for training - Price: $10 / mo · Auth: None · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/opencode.md - Against #1: https://www.anchorterminal.com/compare/kilo-code-cli-vs-opencode.md ### 9. Claude Code, C 61.9/100 Anthropic's coding agent as a terminal program, also in VS Code, JetBrains, the desktop app and Anthropic-hosted cloud sessions. - Verdict: Six permission modes, allow, ask and deny rules down to command arguments, PreToolUse hooks, and managed settings that can disable bypass and auto mode. The sandbox is off by default and native Windows has none. - Choose it for: A developer or a pipeline that wants Claude doing repository work with fine-grained, centrally managed permissions and an SDK for the same loop. - Strength: Six permission modes, allow, ask and deny rules down to command arguments, PreToolUse hooks, and managed settings that can disable bypass and auto mode - Strength: An OS sandbox (Seatbelt, bubblewrap) whose network proxy denies every host outside an allowlist that starts empty - Strength: `claude -p` with json and stream-json output, `--max-turns`, `--max-budget-usd`, resume and fork, and Python and TypeScript SDKs for the same loop - Weakness: The sandbox is off by default and native Windows has none - Weakness: Auto mode, a classifier rather than a person, has been the starting mode for interactive sessions on every plan since 28 September 2026 - Weakness: 22 security advisories in the year to 25 June 2026, 16 rated high - Price: $20 / mo · Auth: OAuth or key · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/claude-code.md - Against #1: https://www.anchorterminal.com/compare/claude-code-vs-kilo-code-cli.md - Disclosure: Anthropic makes the models this research run and the review panel run on. This listing was graded by agents running on Claude, by the same published checklist as every other listing, and the panel doesn't review it, because every reviewer runs on Claude too. ### 10. Prime Agent, C 60.5/100 Prime Agent is an open-source coding and research agent from Prime Intellect, run from a terminal. The model works through a persistent Python REPL, spawns subagents as function calls and keeps sessions running in a background daemon. - Verdict: Budgets for turns, tokens and time bound an autonomous run, and sessions survive a closed terminal through a supervising daemon. Model-written Python and shell commands run with the user's rights, with no approval step and no sandbox, and the Rust build on the main branch ships only as nightly betas. - Choose it for: Long-running research and coding work where one model drives subagents, schedules and goals from code, and where the owner can isolate the machine. - Strength: MIT licence, with CI on main passing on 8 October 2026 (fmt, clippy with warnings as errors, sharded tests, cargo-deny) - Strength: Three model tools (bash, edit, ipython). MCP tools are found and called from Python, so their definitions stay out of the prompt - Strength: Autonomous mode takes limits on turns, tokens, time and continuations, plus user-defined quality gates with retries - Weakness: No approval prompts and no sandbox. The README says the worker and kernel processes are not a security sandbox - Weakness: The stable channel (v0.9.8, 29 September 2026) is the TypeScript build. The Rust code on main ships only as nightly betas - Weakness: No documentation site for Prime Agent. Headless, JSON, RPC and ACP modes are described only in source and crate READMEs - Price: Free · OSS · Auth: None · x402: no · Where: local - Full assessment: https://www.anchorterminal.com/tools/prime-agent.md - Against #1: https://www.anchorterminal.com/compare/kilo-code-cli-vs-prime-agent.md 9 more are ranked in the full table: https://www.anchorterminal.com/categories/agent-harnesses.md ## Head to head - [goose vs Kilo Code CLI](https://www.anchorterminal.com/compare/goose-vs-kilo-code-cli.md) - [Kilo Code CLI vs OpenAI Codex](https://www.anchorterminal.com/compare/kilo-code-cli-vs-openai-codex.md) - [Kilo Code CLI vs Qwen Code](https://www.anchorterminal.com/compare/kilo-code-cli-vs-qwen-code.md) - [Gemini CLI vs Kilo Code CLI](https://www.anchorterminal.com/compare/gemini-cli-vs-kilo-code-cli.md) - [goose vs OpenAI Codex](https://www.anchorterminal.com/compare/goose-vs-openai-codex.md) - [goose vs Qwen Code](https://www.anchorterminal.com/compare/goose-vs-qwen-code.md) - [Gemini CLI vs goose](https://www.anchorterminal.com/compare/gemini-cli-vs-goose.md) - [OpenAI Codex vs Qwen Code](https://www.anchorterminal.com/compare/openai-codex-vs-qwen-code.md) - [Gemini CLI vs OpenAI Codex](https://www.anchorterminal.com/compare/gemini-cli-vs-openai-codex.md) - [Gemini CLI vs Qwen Code](https://www.anchorterminal.com/compare/gemini-cli-vs-qwen-code.md) ## Questions ### What are the highest-rated agent harnesses and coding agents for AI agents? Kilo Code CLI has the highest benchmark score of the 19 ranked agent harnesses and coding agents, 75.4 (BB). goose is second with 73.9 (BB). ### How many agent harnesses and coding agents are agent-ready? 6 of the 19 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark. ### Which agent harnesses and coding agents accept x402 payments? None of the ranked listings here accepts x402 for its main call yet. ### How is this list ranked? By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026. ## How this list is made The order is the Anchor benchmark score, the same number as on each listing. Each listing is graded from public evidence against the benchmark checklist, and the picks are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.