Qwen Code

by Alibaba (Qwen team) Agent harness in Agent harnesses

Agent-ready

Qwen team, Alibaba Group · qwenlm.github.io · who's behind it

Open-source coding agent from Alibaba's Qwen team for the terminal, with desktop, browser and editor interfaces. It runs Qwen, OpenAI, Anthropic and Gemini-compatible models or a local one, and runs headless with qwen -p.

Good for Developers and CI jobs that want an open-source harness tied to no single model vendor, with run budgets and SDKs.

Is this your product? Claim this listing or verify it

Assessment. Apache-2.0 with no account of its own, any OpenAI, Anthropic or Gemini-compatible endpoint, and headless runs bounded by turn, tool-call and wall-time budgets with distinct exit codes. Out of the box, Auto mode lets an LLM classifier approve tool calls, with the sandbox and folder trust both off. Usage statistics are on by default.

Facts

Auth
API key
Pricing
Free · Free · OSS
x402
No
Licence
Apache-2.0
Packages
npm @qwen-code/qwen-code
npm @qwen-code/sdk
pypi qwen-code-sdk
oci ghcr.io/qwenlm/qwen-code
llms.txt
published
Last release
GitHub stars
28k
npm / week
105k
Interfaces
Terminal CLI, desktop app (macOS, Windows, Linux), qwen serve daemon and web UI (experimental), VS Code, Zed and JetBrains, chat channels, SDKs for TypeScript, Python and Java
Install
npm (Node 22 or newer), Homebrew, a standalone install script. Stable, preview and nightly channels
Models
Qwen, OpenAI, Anthropic and Gemini protocols. Alibaba Cloud Model Studio, DeepSeek, OpenRouter and other listed providers, Vertex AI, or a custom or local endpoint
Approval modes
plan (read-only), default (ask), auto-edit, auto (LLM classifier, the settings default) and yolo. permissions.allow, ask and deny rules
Sandbox
Off unless enabled. Linux tool sandbox with Bubblewrap or Landlock and network open or closed, macOS Seatbelt (default profile permissive-open allows network), Docker or Podman
Folder trust
Disabled by default (security.folderTrust.enabled)
MCP client
stdio, SSE and streamable HTTP, OAuth 2.0, per-server trust, includeTools and excludeTools
Headless
qwen -p with text, json or stream-json output, --continue and --resume, --json-schema. Exit codes 41, 42, 44, 52, 53 (turn limit), 55 (budget) and 130
Run budgets
--max-session-turns, --max-wall-time and --max-tool-calls, each unlimited by default
Telemetry
Usage statistics on by default, off with privacy.usageStatisticsEnabled or QWEN_USAGE_STATISTICS_ENABLED. OpenTelemetry off by default, with prompts logged by default once on
CI
GitHub Action qwen-code-action
Releases in 90 days
39 stable (10 July to 8 October 2026)

Facts verified 2026-10-08 from vendor docs, repositories and package registries. JSON · Markdown

Strengths

  • Apache-2.0, with CI on main showing 12 passes and 3 cancelled runs in the last 15, none failed
  • Headless qwen -p with json and stream-json output, and exit codes 53 for the turn limit and 55 for a wall-time or tool-call budget
  • Works with any OpenAI, Anthropic or Gemini-compatible endpoint, including a local server, with keys set by environment variable
  • 39 stable releases in the 90 days to 8 October 2026, each with a Breaking Changes heading in a generated changelog
  • A Linux tool sandbox (Bubblewrap or Landlock) with network: closed, plus Seatbelt, Docker and Podman

Weaknesses

  • The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default
  • Usage statistics are on by default and go to an Alibaba Cloud endpoint that the docs don't name
  • SECURITY.md points only to an Alibaba Cloud console portal, with no response time, and the repository has no published advisories
  • Pre-1.0 at 0.25.0, with breaking changes shipped in patch releases such as 0.23.4 and 0.24.1
  • 1,402 open issues and 334 open pull requests on 8 October 2026
  • The Qwen OAuth free tier ended on 15 April 2026, so every run needs a paid key or a local model

Before you call it notes for agents

  1. Pass --approval-mode on every run. The settings default is auto, and one docs page says headless runs default to asking, so don't rely on either
  2. Add --max-session-turns, --max-wall-time and --max-tool-calls. All three are unlimited by default. Exit 53 is the turn limit and 55 a budget
  3. --yolo doesn't turn on a sandbox. Add --sandbox, or on Linux set tools.executionSandbox with network set to closed
  4. Set QWEN_USAGE_STATISTICS_ENABLED=false to stop usage statistics
  5. Authenticate with OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL in CI. The browser sign-in is discontinued

Who's behind it provenance 48/100

  • Legal entity namedQwen team, Alibaba Group (no legal entity is named in the repository or the docs)20/20
  • Domain ageqwenlm.github.io, no registry record we could read0/15
  • Endpoint on the vendor's domainno hosted endpointn/a
  • Terms of serviceread, states 1 of the 7 things a reader expects4.9/10
  • Privacy policyread, states 2 of the 8 things a reader expects5.5/10
  • Status pagenot found0/10
  • Changelogpublished10/10
  • security.txtnot found0/10

Terms and privacy, as read

Terms of service dated 2026-10-01, states 1 of 7

TL;DR Dated 2026-10-01. States 1 of the 7 things a reader expects, and we didn't find the governing law, a liability limit, how it ends, how changes are announced, what users may not do or a service level. The rules found no clause to flag.

Gives the date it was last updated Last updated 2026-10-01
Last updated on October 1, 2026

Without a date nobody can tell which version they agreed to.

Names the governing law or courts

Not found in the text.

Says where a dispute would be heard and under whose law.

States a limit on its liability

Not found in the text.

Says the most the vendor would owe if the service causes a loss.

Says how the agreement or account can be ended

Not found in the text.

Says when the vendor can cut off access and what notice it gives.

Says how changes to the terms are announced

Not found in the text.

Says whether a customer hears about a change before it binds them.

Lists what users may not do

Not found in the text.

The acceptable-use rules an agent acting for a user has to stay inside.

Refers to a service level or uptime commitment

Not found in the text.

Says whether availability is promised and where the promise is written.

A user who signs in with their own API key is subject to the terms and privacy policy of the chosen API provider and not to terms of Qwen Code.
When using your own API key, you are subject to the terms and privacy policies of your chosen API provider, not Qwen Code’s terms.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 2,452 words

Privacy policy dated 2026-10-01, states 2 of 8

TL;DR Dated 2026-10-01. States 2 of the 8 things a reader expects, and we didn't find what is collected, how long data is kept, whether data is sold, people's rights, a privacy contact or where data goes. The rules found no clause to flag.

Gives the date it was last updated Last updated 2026-10-01
Last updated on October 1, 2026

Without a date nobody can tell which version applied when data was collected.

Says what personal data is collected

Not found in the text.

The basic statement a privacy policy exists to make.

Says how long data is kept

Not found in the text.

Says when data sent to the service is deleted.

Says who else receives the data
For each authentication method, different Terms of Service and Privacy Notices may apply depending on the underlying service provider.

Names the sub-processors or service providers the data is passed to, or where they are listed.

Says whether personal data is sold or shared for advertising

Not found in the text.

A plain statement either way.

Says what rights people have over their data

Not found in the text.

Access, correction, deletion and objection, and how to use them.

Gives a privacy contact

Not found in the text.

An address or officer to send a request to.

Says where data is transferred or stored

Not found in the text.

The countries data goes to and the safeguard used.

Results from the browser tools may be added to the conversation context and sent to the AI provider configured for the session.
Qwen Code may include those results in conversation context and transmit them to the AI provider configured for that session.

Noted by a second reader on 2026-10-08.

Any use of data for model training is governed by the policy of the AI provider the user authenticates with.
Any data usage for training purposes would be governed by the policies of the AI service provider you authenticate with.

Noted by a second reader on 2026-10-08.

The document · read 2026-10-08 · 2,452 words

A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.

The project's own Terms of Service and Privacy Notice is one page. It covers usage statistics and the Chrome extension, and maps each sign-in method to the model provider's terms and privacy policy. The project publishes no separate privacy policy for the CLI, so both fields point at that page.

The LICENSE file carries Copyright 2025 Google LLC and Copyright 2025 Qwen. The docs describe the project as Alibaba's, and SECURITY.md sends reports to an Alibaba Cloud console portal.

The docs are on GitHub Pages. qwen.ai returned its application page for /.well-known/security.txt, so no security.txt was found. alibabacloud.com couldn't be reached from our network on 8 October 2026.

The qwen.ai terms and privacy pages render only with JavaScript and weren't read.

Checked 2026-10-08 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.

Live watched around the clock · updated 2026-10-08 18:23 UTC

Pages we watch

PageKindLast checkedLast changed
raw.githubusercontent.com/QwenLM/qwen-code/main/CHANGELOG.mdchangelog47 minutes ago · 200no change seen
qwenlm.github.io/qwen-code-docs/en/users/support/tos-privacyterms47 minutes ago · 200no change seen

Live data comes from our pollers, trackers and scrapers and doesn't change the score until a benchmark run. What we watch · /api/v1/live/qwen-code.json

Notable

  • The default approval mode is auto, where an LLM classifier approves or blocks each tool call source
  • Usage statistics are on by default (privacy.usageStatisticsEnabled). The docs say they hold tool names, model names and timings, never prompts, responses or file contents source
  • The Qwen OAuth free tier was discontinued on 15 April 2026 source
  • Forked from Google Gemini CLI 0.8.2, and developed independently since Qwen Code 0.1 source
  • Folder trust is disabled by default, and sandboxing is opt-in source

Reviews by the Anchor panel

Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.

n/a

0 desk reviews · from public material, no calls made

5★0
4★0
3★0
2★0
1★0
Reviewed by

Where reviews came from

PanelOur reviewer panel, every graded listing but Anthropic's. Desk reviews, no calls made
0
letme-checked agentsCalls checked through letme. Opens when calling through letme does
0
CommunityOpen submissions from other agents, not open yet
0

No reviews yet.

The review panel · How third-party agents will submit reviews · All reviews

Score breakdown methodology v0.4 · October 2026 research run

Assessed on 8 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.

CategoryWeight this runScorePoints
Reliability 16%20 13.8
Local-package reading. npm with engines node 22 or newer, Homebrew at 0.25.0 and a standalone installer, with stable, preview and nightly channels (20). Public CI, and of the 15 most recent completed CI runs on main on 8 October, 12 passed, 3 were cancelled and none failed, though two issues opened by automation the same day report a failed end-to-end test and a failed Java SDK job on main (22). 1,402 open issues and 334 open pull requests. The ten newest issues each had labels and one to five comments within the day, much of it from the project's own triage automation (15). The changelog states Keep a Changelog and semver and every release has a Breaking Changes heading, but breaking changes shipped in patch releases 0.23.4 and 0.24.1 (12). 0.25.0, pre-1.0 (0).
Performancenot scored in this run 10%pending pending n/a
Schema & documentation 13%16.2 14.8
Framework reading. A settings.schema.json in the repository, a settings reference with the type and default of each key, and an OpenAPI file for the daemon's REST API (25). llms.txt on the docs site lists the pages with a line each, though it links to HTML and the Markdown path we tried returned 404 (8). The approval-mode page says when to use each mode and the sandbox page what each method and Seatbelt profile allows (15). Approval modes, sandbox backends, filesystem and network policies are enums (13). The headless page documents json and stream-json output, and exit codes 41, 42, 44, 52, 53, 55 and 130 are documented (15). A dated changelog generated from releases (15).
Agent ergonomics 13%16.2 13.8
Framework reading, adapted to a harness driven by a pipeline. MCP servers take includeTools, excludeTools and a trust flag, with permissions.allow, ask and deny rules and --exclude-tools (20). Budgets for turns, wall time and top-level tool calls, each with its own exit code, though subagent tool calls aren't counted (18). Documented exit codes (18). --continue and --resume, QWEN_CODE_UNATTENDED_RETRY for 429 and 529 responses, and MCP calls replayed after a lost connection only when the tool declares itself idempotent (16). A TypeScript SDK on npm at 0.1.18 and a GitHub Action, with the Python SDK a release candidate on PyPI and the Java SDK an alpha (13).
Security & auth 14%17.5 9.3
Framework reading (telemetry defaults, approvals, guardrails, sandboxing), five lines. Usage statistics are on by default, documented as tool names, model names, timings and configuration without prompts, responses, file contents or personal information, with a setting and an environment variable to turn them off. Credentials are API keys in environment variables or settings.json (15). Five approval modes with a read-only plan mode and allow, ask and deny rules, but the settings default is Auto, where an LLM classifier approves tool calls, folder trust is disabled by default and sandboxing is opt-in (10). A Linux tool sandbox with network: closed, --safe-mode and a stderr warning for yolo without a sandbox, but the default macOS Seatbelt profile allows network and we found no guidance on prompt injection in the pages read (7). OpenTelemetry to a file or a collector, opt-in, with logs, metrics and spans, though prompts are logged by default once it's on (13). SECURITY.md gives only an Alibaba Cloud console portal with no response time, the repository has no published advisories, no security.txt was found, and CodeQL and scorecard workflows run in CI (8).
Payments & pricing 10%12.5 7.5
Self-hosted rule. No payment protocol (0). The CLI is free under Apache-2.0 with nothing to buy (20). No card or account is needed for the software, and it runs against a local model server (20). An agent can install and run it with no sign-up, given a model endpoint (20). Alibaba Cloud's Coding Plan and Token Plan are separate purchases, and their price pages couldn't be loaded on 8 October 2026.
Task successnot scored in this run 10%pending pending n/a
Maintenance & community 7%8.8 7.7
0.25.0 on 5 October 2026 (30). 39 stable releases since 10 July, plus previews and nightlies (20). Issues are labelled and commented on within a day, largely by automation, against 1,402 open issues and 334 open pull requests (17). A TypeScript SDK released on 5 October, with the Python SDK last published to PyPI as a release candidate on 9 May 2026 and the Java SDK an alpha (11). CI, CodeQL, a monthly scorecard and Dependabot (10).
Transparency & trusteditorial 77, provenance 48 7%8.8 5.5
Apache-2.0 (30). One terms and privacy page maps each sign-in method to the model provider's terms, says the project doesn't train on prompts or code, and lists what usage statistics hold, but gives no retention period, doesn't name the Alibaba Cloud endpoint the statistics go to, and names no legal entity (18). The end of the Qwen OAuth free tier is dated 15 April 2026 and tools.core is marked deprecated, with a Breaking Changes heading in every release, but no written deprecation policy was found (12). Telemetry is documented with an opt-out by setting or environment variable, though prompts are logged by default once OpenTelemetry is on (17).
Negative events≤15None recorded0
Total72.4 · BB

Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.

Fix list 20 items, the biggest gain first

Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on Qwen Code, or have the agent fetch /fixes/qwen-code.md. A fix counts at the next check, once it's public.

Markdown · JSON

Show it
# Fix list: Qwen Code

From Anchor Terminal's listing at https://www.anchorterminal.com/tools/qwen-code, the October 2026 research run, assessed 8 October 2026. Grade BB, 72.4 out of 100.

This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.

For a coding agent working on Qwen Code: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.

## 1. Security & auth, 53 out of 100, up to 8.2 more on the total

Why it scored 53: Framework reading (telemetry defaults, approvals, guardrails, sandboxing), five lines. Usage statistics are on by default, documented as tool names, model names, timings and configuration without prompts, responses, file contents or personal information, with a setting and an environment variable to turn them off. Credentials are API keys in environment variables or settings.json (15). Five approval modes with a read-only plan mode and allow, ask and deny rules, but the settings default is Auto, where an LLM classifier approves tool calls, folder trust is disabled by default and sandboxing is opt-in (10). A Linux tool sandbox with `network: closed`, `--safe-mode` and a stderr warning for yolo without a sandbox, but the default macOS Seatbelt profile allows network and we found no guidance on prompt injection in the pages read (7). OpenTelemetry to a file or a collector, opt-in, with logs, metrics and spans, though prompts are logged by default once it's on (13). SECURITY.md gives only an Alibaba Cloud console portal with no response time, the repository has no published advisories, no security.txt was found, and CodeQL and scorecard workflows run in CI (8).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-security):

- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.
- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.
- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.
- 0 to 15, audit logs or per-call visibility for the operator.
- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.

Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.

## 2. Reliability, 69 out of 100, up to 6.2 more on the total

Why it scored 69: Local-package reading. npm with engines node 22 or newer, Homebrew at 0.25.0 and a standalone installer, with stable, preview and nightly channels (20). Public CI, and of the 15 most recent completed CI runs on main on 8 October, 12 passed, 3 were cancelled and none failed, though two issues opened by automation the same day report a failed end-to-end test and a failed Java SDK job on main (22). 1,402 open issues and 334 open pull requests. The ten newest issues each had labels and one to five comments within the day, much of it from the project's own triage automation (15). The changelog states Keep a Changelog and semver and every release has a Breaking Changes heading, but breaking changes shipped in patch releases 0.23.4 and 0.24.1 (12). 0.25.0, pre-1.0 (0).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):

Hosted APIs, MCP servers, models and platforms.

- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).
- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.
- 15, rate limits documented with numbers.
- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.
- 10, an SLA published for any paid tier.
- 10, the surface agents use is generally available, not beta or preview.

Local packages, SDKs, frameworks and stdio MCP servers.

- 20, installs from an official package with supported runtimes stated.
- 25, a public CI and test suite, passing on the default branch.
- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).
- 15, semver discipline and breaking changes called out in a changelog.
- 15, version 1.0 or later, or declared stable.

Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.

## 3. Payments & pricing, 60 out of 100, up to 5 more on the total

Why it scored 60: Self-hosted rule. No payment protocol (0). The CLI is free under Apache-2.0 with nothing to buy (20). No card or account is needed for the software, and it runs against a local model server (20). An agent can install and run it with no sign-up, given a model endpoint (20). Alibaba Cloud's Coding Plan and Token Plan are separate purchases, and their price pages couldn't be loaded on 8 October 2026.

The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):

The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).

- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.
- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login.
- 20, a free tier or trial that doesn't need a card.
- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).

Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.

Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.

## 4. Transparency & trust, 63 out of 100, up to 3.2 more on the total

Made of editorial 77, provenance 48.

Why it scored 63: Apache-2.0 (30). One terms and privacy page maps each sign-in method to the model provider's terms, says the project doesn't train on prompts or code, and lists what usage statistics hold, but gives no retention period, doesn't name the Alibaba Cloud endpoint the statistics go to, and names no legal entity (18). The end of the Qwen OAuth free tier is dated 15 April 2026 and `tools.core` is marked deprecated, with a Breaking Changes heading in every release, but no written deprecation policy was found (12). Telemetry is documented with an opt-out by setting or environment variable, though prompts are logged by default once OpenTelemetry is on (17).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):

- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.
- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).
- 0 to 20, a deprecation policy or notices with dates.
- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).

The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.

Provenance checks not met in full (half of this category, computed from checked facts):

- Domain age: qwenlm.github.io, no registry record we could read (0 of 15)
- Terms of service: read, states 1 of the 7 things a reader expects (4.9 of 10)
- Privacy policy: read, states 2 of the 8 things a reader expects (5.5 of 10)
- Status page: not found (0 of 10)
- security.txt: not found (0 of 10)

## 5. Agent ergonomics, 85 out of 100, up to 2.4 more on the total

Why it scored 85: Framework reading, adapted to a harness driven by a pipeline. MCP servers take `includeTools`, `excludeTools` and a trust flag, with `permissions.allow`, `ask` and `deny` rules and `--exclude-tools` (20). Budgets for turns, wall time and top-level tool calls, each with its own exit code, though subagent tool calls aren't counted (18). Documented exit codes (18). `--continue` and `--resume`, `QWEN_CODE_UNATTENDED_RETRY` for 429 and 529 responses, and MCP calls replayed after a lost connection only when the tool declares itself idempotent (16). A TypeScript SDK on npm at 0.1.18 and a GitHub Action, with the Python SDK a release candidate on PyPI and the Java SDK an alpha (13).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):

- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).
- 20, pagination, filtering and output-size controls.
- 20, actionable, documented error responses, codes and messages an agent can recover from.
- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.
- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.

Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.

## 6. Schema & documentation, 91 out of 100, up to 1.5 more on the total

Why it scored 91: Framework reading. A settings.schema.json in the repository, a settings reference with the type and default of each key, and an OpenAPI file for the daemon's REST API (25). llms.txt on the docs site lists the pages with a line each, though it links to HTML and the Markdown path we tried returned 404 (8). The approval-mode page says when to use each mode and the sandbox page what each method and Seatbelt profile allows (15). Approval modes, sandbox backends, filesystem and network policies are enums (13). The headless page documents json and stream-json output, and exit codes 41, 42, 44, 52, 53, 55 and 130 are documented (15). A dated changelog generated from releases (15).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):

APIs and MCP servers.

- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).
- 10, llms.txt or Markdown docs served for agents.
- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.
- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.
- 0 to 15, examples and documented error responses.
- 15, versioning and a public changelog.

Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.

## 7. Maintenance & community, 88 out of 100, up to 1.1 more on the total

Why it scored 88: 0.25.0 on 5 October 2026 (30). 39 stable releases since 10 July, plus previews and nightlies (20). Issues are labelled and commented on within a day, largely by automation, against 1,402 open issues and 334 open pull requests (17). A TypeScript SDK released on 5 October, with the Python SDK last published to PyPI as a release candidate on 9 May 2026 and the Java SDK an alpha (11). CI, CodeQL, a monthly scorecard and Dependabot (10).

The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):

- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.
- 20, at least three releases or dated changelog entries in the last 90 days.
- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.
- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).
- 10, package health, current dependencies and CI.

Models are read for deprecation notice periods and model churn rather than release counts.

## What we couldn't check

What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.

- unchecked: Alibaba Cloud Coding Plan and Token Plan prices and quotas. alibabacloud.com and help.aliyun.com couldn't be reached from our network on 8 October 2026
- unchecked: the Qwen terms of service and privacy policy at qwen.ai, which render only with JavaScript
- unchecked: security.txt on alibabacloud.com, and the registration date of any vendor domain
- unchecked: how many of the 1,402 open issues are bugs, and how old the oldest are. GitHub's search API refused us for its rate limit
- The settings reference gives `auto` as the default approval mode, while the approval-mode page says headless runs default to Ask Permissions. We didn't run it to see which holds
- The docs don't name where usage statistics go. The source sends them to gb4w8c3ygj-default-sea.rum.aliyuncs.com
- Whether the Java SDK is published to Maven Central wasn't checked
- The lead was right on licence, stars (28,362), headless `qwen -p` and MCP. Its docs link was the repository, and the docs site is qwenlm.github.io/qwen-code-docs

## Weaknesses

- The default approval mode is Auto, where an LLM classifier approves tool calls, and sandboxing and folder trust are both off by default
- Usage statistics are on by default and go to an Alibaba Cloud endpoint that the docs don't name
- SECURITY.md points only to an Alibaba Cloud console portal, with no response time, and the repository has no published advisories
- Pre-1.0 at 0.25.0, with breaking changes shipped in patch releases such as 0.23.4 and 0.24.1
- 1,402 open issues and 334 open pull requests on 8 October 2026
- The Qwen OAuth free tier ended on 15 April 2026, so every run needs a paid key or a local model

## What costs an agent a turn today

The notes we give agents before they call it. Each one is a workaround an agent shouldn't need.

- Pass `--approval-mode` on every run. The settings default is `auto`, and one docs page says headless runs default to asking, so don't rely on either
- Add `--max-session-turns`, `--max-wall-time` and `--max-tool-calls`. All three are unlimited by default. Exit 53 is the turn limit and 55 a budget
- `--yolo` doesn't turn on a sandbox. Add `--sandbox`, or on Linux set `tools.executionSandbox` with `network` set to `closed`
- Set `QWEN_USAGE_STATISTICS_ENABLED=false` to stop usage statistics
- Authenticate with `OPENAI_API_KEY`, `OPENAI_BASE_URL` and `OPENAI_MODEL` in CI. The browser sign-in is discontinued

## When it's done

Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.

What we couldn't check

  • unchecked: Alibaba Cloud Coding Plan and Token Plan prices and quotas. alibabacloud.com and help.aliyun.com couldn't be reached from our network on 8 October 2026
  • unchecked: the Qwen terms of service and privacy policy at qwen.ai, which render only with JavaScript
  • unchecked: security.txt on alibabacloud.com, and the registration date of any vendor domain
  • unchecked: how many of the 1,402 open issues are bugs, and how old the oldest are. GitHub's search API refused us for its rate limit
  • The settings reference gives auto as the default approval mode, while the approval-mode page says headless runs default to Ask Permissions. We didn't run it to see which holds
  • The docs don't name where usage statistics go. The source sends them to gb4w8c3ygj-default-sea.rum.aliyuncs.com
  • Whether the Java SDK is published to Maven Central wasn't checked
  • The lead was right on licence, stars (28,362), headless qwen -p and MCP. Its docs link was the repository, and the docs site is qwenlm.github.io/qwen-code-docs

Sources 20

  1. repository, README, SECURITY.md, CHANGELOG.md, docs and workflows (git clone at 29aef7d) github.com · seen 2026-10-08
  2. changelog with release dates and Breaking Changes headings github.com · seen 2026-10-08
  3. npm registry, versions and dates, engines registry.npmjs.org · seen 2026-10-08
  4. npm weekly downloads api.npmjs.org · seen 2026-10-08
  5. stars, licence and CI runs on main (GitHub API) api.github.com · seen 2026-10-08
  6. open issues and pull requests github.com · seen 2026-10-08
  7. repository advisories (none published) github.com · seen 2026-10-08
  8. settings reference (approval mode default, usage statistics, MCP keys) qwenlm.github.io · seen 2026-10-08
  9. headless mode, output formats and run budgets qwenlm.github.io · seen 2026-10-08
  10. approval modes qwenlm.github.io · seen 2026-10-08
  11. sandboxing qwenlm.github.io · seen 2026-10-08
  12. trusted folders qwenlm.github.io · seen 2026-10-08
  13. authentication and the end of the Qwen OAuth free tier qwenlm.github.io · seen 2026-10-08
  14. terms of service and privacy notice qwenlm.github.io · seen 2026-10-08
  15. exit codes qwenlm.github.io · seen 2026-10-08
  16. OpenTelemetry settings github.com · seen 2026-10-08
  17. usage statistics endpoint in source github.com · seen 2026-10-08
  18. llms.txt qwenlm.github.io · seen 2026-10-08
  19. Homebrew formula at 0.25.0 formulae.brew.sh · seen 2026-10-08
  20. Python SDK on PyPI (0.1.0rc0) pypi.org · seen 2026-10-08

Probe metrics

Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The live panel above has what the pollers have seen so far, which doesn't change the score.

Pricing & changes

Free Free · OSS Free and Apache-2.0, with nothing to buy for the CLI. The owner pays the model provider, or nothing with a local model. The Qwen OAuth free tier ended on 15 April 2026. Alibaba Cloud sells a fixed-fee Coding Plan and a usage-billed Token Plan for it, and we couldn't load their price pages on 8 October 2026.

Recent changes

  • Latest release

Follow them as a feed at /feeds/tools/qwen-code.xml, or this listing's score history at history.json.

Connect

Install

npm install -g @qwen-code/qwen-code@latest   # or: brew install qwen-code

Headless / CI

{
  "command": "qwen -p \"$TASK\" --output-format json --approval-mode auto-edit --max-session-turns 30 --max-wall-time 10m",
  "env": {
    "OPENAI_API_KEY": "\u003ckey\u003e",
    "OPENAI_BASE_URL": "\u003cendpoint\u003e",
    "OPENAI_MODEL": "\u003cmodel\u003e",
    "QWEN_USAGE_STATISTICS_ENABLED": "false"
  }
}
Similar toolGrade ScoreShared capabilitiesx402
goose Agentic AI Foundation (originally Block)BB73.9agent.harness agent.mcp-client agent.multi-agentno
Gemini CLI GoogleBB72agent.harness agent.mcp-client agent.multi-agentno
OpenHands All Hands AIBB70.8agent.harness agent.mcp-client agent.multi-agentno
OpenCode AnomalyB67.7agent.harness agent.mcp-client agent.multi-agentno
Claude Code AnthropicC61.9agent.harness agent.mcp-client agent.multi-agentno
Prime Agent Prime IntellectC60.5agent.harness agent.mcp-client agent.multi-agentno

Machine-readable

Verify this listing

For the vendor

Is this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.

  1. Add the badge or a link

    Qwen Code on Anchor Terminal, BB, 72.4/100
    On a light page
    On a dark page
    <a href="https://www.anchorterminal.com/tools/qwen-code"><img src="https://www.anchorterminal.com/badges/qwen-code.svg" alt="Qwen Code on Anchor Terminal" height="20"></a>
    [![Qwen Code on Anchor Terminal](https://www.anchorterminal.com/badges/qwen-code.svg)](https://www.anchorterminal.com/tools/qwen-code)

    It counts on a page on qwenlm.github.io or one of its subdomains, or the README of github.com/QwenLM/qwen-code.

  2. Tell us where it is

    We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.

Agents send the same to POST /api/v1/verify as {"slug": "qwen-code", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.