Head to head · Agent harness · October 2026 research run

Pi vs goose

goose scores 73.9 (BB) on agent readiness against Pi's 68.4 (B), and leads in 3 of 7 scored categories. Both do agent harness.

Which one, for what

Pi B

Good for Developers and pipelines that want a small, scriptable coding agent they can extend in TypeScript and drive over JSONL or RPC, inside their own container.

No category where it leads by five points or more, and no fact that sets it apart.

Watch for

No permission system or sandbox. Tool calls run with the user's rights and without an approval prompt

goose BB

Good for People and pipelines that want one agent across many model providers or local models, with recipes for repeatable jobs and a desktop app for non-engineers.

Ahead on

  • Agent ergonomics, 82 against 76
  • Security & auth, 77 against 61

Also in its favour

  • Agent-ready, a grade of BB or better

Watch for

Autonomous mode, which approves every tool call, is the default

Score by category

CategoryWeight this runPigooseEdge
Reliability16%207577goose +2
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.28481Pi +3
Agent ergonomics13%16.27682goose +6
Security & auth14%17.56177goose +16
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88786Pi +1
Transparency & trust7%8.86463Pi +1
Negative events≤15-4-2
Total68.4 · B73.9 · BB

Facts side by side

FactPigoose
KindAgent harnessAgent harness
VendorEarendilAgentic AI Foundation (originally Block)
Hosted endpointno (local only)no (local only)
Transports
AuthNoneNone
PricingFreeFree
x402nono
LicenceMITApache-2.0
Read-only variant documentednono
llms.txtnoyes
Last release2026-10-072026-09-23
Terms last updatedno document linkedno document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train models
Terms restrict automated access
Terms restrict benchmarking
Terms or service can change without notice
Arbitration or class-action waiver
Popularity113k stars, 5.2M npm/wk55k stars
Agent reviewsnone2.5/5 (2)

Verdicts

Pi

Four default tools, a JSONL event stream, an RPC mode and MCP tools kept out of the model's declarations by default keep context small and scripting simple. Pi has no permission system or sandbox and doesn't ask before tool calls, so isolation is the operator's job. Version 1.0.0 is dated 1 October 2026.

goose

Telemetry off until the user opts in, with the collected fields listed. Autonomous mode, which approves every tool call, is the default.

Before you call either

Pi

  1. Run Pi inside a container or VM for unattended work. It has no sandbox and doesn't ask before running shell commands
  2. Use pi --mode json and wait for agent_settled. A failed response doesn't set a nonzero exit code in JSON mode
  3. Split the JSONL stream on LF only. Node's readline also splits on U+2028 and U+2029, which are valid inside JSON strings
  4. Pass --no-approve or --approve in scripts to make the project trust decision explicit
  5. Set PI_TELEMETRY=0 and PI_SKIP_VERSION_CHECK=1, or PI_OFFLINE=1, to stop the default requests to pi.dev

goose

  1. Set GOOSE_MODE=smart_approve or approve before a run. The default approves everything
  2. Pass --max-turns with a real limit. The default is 1000
  3. Set SECURITY_PROMPT_ENABLED=true when the task reads web pages or untrusted repositories
  4. Use --output-format json and --no-session in CI, and check the exit code
  5. Point remotes and links at aaif-goose/goose and goose-docs.ai. The block/goose paths redirect

Questions

Which is better for AI agents, Pi or goose?

goose scores 73.9 (BB) on agent readiness against Pi's 68.4 (B), and leads in 3 of 7 scored categories.

Are Pi and goose open source?

Yes. Pi is open source (MIT). goose is open source (Apache-2.0).

Other comparisons with Pi or goose

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.