# Fix list: Agno From Anchor Terminal's listing at https://www.anchorterminal.com/tools/agno, the October 2026 research run, assessed 8 October 2026. Grade BB, 72.8 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Agno: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Security & auth, 64 out of 100, up to 6.3 more on the total Why it scored 64: Telemetry is on by default. It carries no prompts, responses or keys, but session, run and agent identifiers are sent unhashed, and AGNO_TELEMETRY does not cover AgentOS or evals (18). Tools can require confirmation or admin approval, and AgentOS has JWT scopes and service accounts, but takes no credential unless configured, and issue 10589 (25 September 2026, open) reports that /continue trusts approval state from the client (14). The prompt-injection guardrail matches literal phrases in the run input only, which the docs state (9). OpenTelemetry tracing, run history and an audit log in the 3.1 authorisation package (13). A valid security.txt pointing to GitHub advisories and three advisories published in the repository. No SECURITY.md, no bounty found, and the trust centre could not be read (10). Framework reading. The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 2. Payments & pricing, 50 out of 100, up to 6.3 more on the total Why it scored 50: No payment protocol (0). Self-hosted rule, with the paid option scored. The control plane has public plan prices, Free at $0 and Pro at $150 a month, with Enterprise by quote (10). The package installs with no card (20) and runs with no Agno account (20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 3. Agent ergonomics, 76 out of 100, up to 3.9 more on the total Why it scored 76: An agent with an MCP server is about ten lines through MCPTools, with include_tools, exclude_tools and tool_name_prefix (22). tool_call_limit exists but defaults to None, and sessions, history and compression are configurable (14). A typed exception hierarchy, with RetryAgentRun and StopAgentRun to steer the loop from a tool (16). Runs continue by run ID after a pause and AgentOS has a durable queue, which is off by default (16). Few required parameters, Python only (8). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 4. Reliability, 83 out of 100, up to 3.4 more on the total Why it scored 83: Official PyPI package with Python 3.9 or later stated (20). Public CI runs Ruff, mypy and the unit tests in five groups, and of the 12 latest runs on main, 10 passed and 2 were cancelled (25). 867 open issues, 171 of them labelled bug and 227 older than six months, with 287 opened and 105 closed in the 30 days to 8 October. Recent bug reports mostly have one to three comments (12). Majors 2.0 and 3.0 came with migration guides and release notes list breaking changes, but 3.1.0, a minor, shipped a breaking filesystem table re-key (11). Version 3.1.1 with the Production/Stable classifier (15). Local-software reading. The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 5. Schema & documentation, 90 out of 100, up to 1.6 more on the total Why it scored 90: A class reference under docs.agno.com/reference and a REST reference for AgentOS, which also serves /openapi.json (25). llms.txt with 3,963 lines, llms-full.txt, a Markdown twin of each page and a docs MCP server (10). Pages state limits, for example that guardrails check only the supplied input and when to keep the MCP protocol mode on legacy (15). Pydantic-typed classes and tool signatures, with a py.typed marker (13). Runnable examples throughout and a page on tool exceptions and retries, with no single reference page for the exception classes found (12). Dated changelog and release notes per version (15). Framework reading. The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 6. Transparency & trust, 83 out of 100, up to 1.5 more on the total Made of editorial 73, provenance 93. Why it scored 83: Apache 2.0 (30). The telemetry page lists the fields sent and says identifiers are not hashed, and the pricing page says the control plane holds no copy of customer data. The terms and privacy pages render only with JavaScript and were not read, and no retention period for telemetry was found (15). Migration guides for 2.0 and 3.0 and migration notices on what was removed, with no written deprecation policy found (10). Telemetry is disclosed in the README and docs with an opt-out, though the environment variable does not cover AgentOS or evals (18). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Terms of service: published, but our reader couldn't read it (7 of 10) - Privacy policy: published, but our reader couldn't read it (7 of 10) ## 7. Maintenance & community, 86 out of 100, up to 1.2 more on the total Why it scored 86: 3.1.1 on 2 October 2026 (30). 25 stable releases and 5 pre-releases on PyPI in the 90 days to 8 October (20). Bug reports are usually answered within days, but 867 issues and 968 pull requests are open, and issues were opened nearly three times as fast as they were closed in the last 30 days (14). The Python package is current (15). CI runs lint, types and tests with pinned development tools. No Dependabot configuration file is in the repository (7). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## Deductions Each comes off the total. A fixed and documented problem counts for less at the next check. - 2026-06-30. GHSA-m4w4-3r8h-jw98 (CVE-2026-35002, critical), arbitrary code execution through eval() on field_type in a FunctionCall, fixed in 2.3.24. GHSA-w9j5-7p53-pvr3 (CVE-2026-10105, high, published 2026-07-27), SQL injection in the ClickHouse vector store's delete_by_metadata, parameterised in the current source. GHSA-vw84-hprm-cxmm (CVE-2025-64168, high, 2025-10-31), session state written to the wrong session under concurrency, fixed in 2.2.2. All three are fixed and published as advisories in the repository, so 1 point each. https://github.com/agno-agi/agno/security/advisories ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: the terms of service and privacy policy at os.agno.com/legal/tos and /legal/privacy. Both pages render only with JavaScript and their text was not read - unchecked: trust.agno.com, which returned only a heading without JavaScript, so any SOC 2 or ISO 27001 status is unknown - Whether the control plane's free plan asks for a card. The pricing page does not say - How long telemetry events are kept. The telemetry page gives no retention period - OSV dates the eval injection and SQL injection records 2026-04-02 and 2026-05-29, earlier than the repository advisories of 2026-06-30 and 2026-07-27. The first was reported through VulnCheck - Whether issue 10589 (approval state trusted from the client in /continue) is confirmed by the maintainers. It was open with two comments on 8 October 2026 - The lead held (Python SDK, AgentOS with REST and MCP, formerly Phidata). The category was kept as frameworks ## Weaknesses - Usage telemetry is on by default and sends session, run and agent identifiers unhashed. Prompts and outputs are not sent - AGNO_TELEMETRY=false does not cover AgentOS launches or evals, which need telemetry=False on the instance - Three advisories in 12 months, an eval injection rated critical, a ClickHouse SQL injection and a session state leak, all fixed - 867 open issues and 968 open pull requests on 8 October 2026, with 287 issues opened and 105 closed in 30 days - AgentOS requires no credential unless a JWT key or OS_SECURITY_KEY is configured - Terms and privacy pages at os.agno.com render only with JavaScript and could not be read ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Set AGNO_TELEMETRY=false before the first run, and pass telemetry=False to AgentOS and to each eval, which ignore the variable - Configure JWT_VERIFICATION_KEY with AgentOS authorisation switched on, or OS_SECURITY_KEY, before exposing AgentOS. With neither set, REST and MCP routes take no credential - Mark tools that write with requires_confirmation, and resume with continue_run once every requirement is resolved - Set tool_call_limit on agents. The default is None - Install extras for what you use, such as agno[mcp] for MCPTools and agno[os,mcp] for AgentOS with its MCP server ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.