Best of · Agent frameworks
Best agent frameworks and SDKs
All 10 ranked agent frameworks and SDKs on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.
- 10 ranked
- 9 agent-ready
- Updated 8 October 2026
Top three
Picks by need
Worked out from the scores, prices and facts, so they change when the research does.
Schema & documentation
100/100 on schema & documentation, against 95 for the overall leader.
The shortlist
| # | Tool | Grade | Best for | Price | Where |
|---|---|---|---|---|---|
| 1 | OpenAI Agents SDK OpenAI |
AA 86.2 | Teams that want hand-offs, guardrails and approvals with little code, and are happy with OpenAI's tracing or will turn it off. | Free · OSS | library |
| 2 | Pydantic AI Pydantic |
A 83.7 | Python teams that want typed, validated outputs and tools, many model providers and durable runs on an engine they already use. | Free · OSS | library |
| 3 | Microsoft Agent Framework Microsoft |
A 82.2 | Teams on Azure or .NET, and for teams moving off AutoGen or Semantic Kernel, that want MCP, approval, checkpointed workflows and OpenTelemetry in one SDK. | Free · OSS | library |
| 4 | Docker Agent Docker |
BB 76.5 | Teams already on Docker who want agents defined in a file, shared through an OCI registry and run the same way in a terminal, in CI or behind an HTTP, MCP or A2A server. | Free · OSS | library |
| 5 | AgentOS Framers Lab, Inc. |
BB 75.3 | TypeScript teams that want long-term memory, personas, voice and channel adapters and multi-agent teams in one package, and that will pin versions. | Free · OSS | library |
| 6 | Agent Development Kit (ADK) |
BB 74.7 | Teams on Google Cloud or working outside Python, and for anyone who wants approval, guardrail callbacks and evaluation in one framework. | Free · OSS | library |
| 7 | Agno Agno Inc. |
BB 72.8 | Python teams that want agents, teams and workflows plus a self-hosted server with REST, MCP, approvals and scheduling in one package. | Free · OSS | library |
| 8 | Claude Agent SDK Anthropic |
BB 72.1 | Agents that work on files, code and shells and want Claude Code's tool set and permission model without building them. | Free | library |
| 9 | LangGraph LangChain |
BB 70.6 | Long-running, stateful agents that pause for people and must survive restarts. | Free · OSS | library |
| 10 | CrewAI CrewAI |
B 67 | Multi-agent role-and-task workflows and for teams that want Flows for ordered, stateful steps. | Free · OSS | library |
How to choose
- Telemetry sent by defaultCheck what the framework sends by default and how to switch it off, since traces can carry prompts, tool arguments and customer data to a third party.
- Durable state and resumptionCheck where run state is stored and whether a run resumes after a crash mid-tool-call, because an agent that restarts from scratch can repeat payments or emails.
- Human approval enforcementCheck whether approval gates run in code or only appear in prompts, since a model can ignore a prompt rule but a code check is harder to get past.
- Multi-agent hand-offs and tracingCheck how hand-offs pass context and whether each step appears in traces, since a lost hand-off is hard to debug without a record of what each agent saw.
Each one in detail
OpenAI Agents SDK
AA 86.2/100Multi-agent framework built on agents, hand-offs and guardrails, with sessions, tracing and human approval.
Verdict MCP in about 11 lines, with static and dynamic tool filters and require_approval. Tracing on by default, with model and tool content, sent to OpenAI.
Choose it for Teams that want hand-offs, guardrails and approvals with little code, and are happy with OpenAI's tracing or will turn it off.
Strengths
- MCP in about 11 lines, with static and dynamic tool filters and require_approval
- Input and output guardrails, built-in human approval and sandbox agents
- Typed exceptions plus error_handlers, and RunState to resume a paused run
Weaknesses
- Tracing on by default, with model and tool content, sent to OpenAI
- Tracing isn't available to zero-data-retention organisations
- Pre-1.0, and minor releases can carry breaking changes, as 0.20.0, 0.21.0 and 0.22.0 did
Price Free · OSSAuth API keyx402 nolibrary
Pydantic AI
A 83.7/100Typed Python agent framework for 25+ model providers, with MCP, A2A and durable execution.
Verdict Typed outputs and tools, validated by Pydantic, with failed validations sent back to the model. Python only.
Choose it for Python teams that want typed, validated outputs and tools, many model providers and durable runs on an engine they already use.
Strengths
- Typed outputs and tools, validated by Pydantic, with failed validations sent back to the model
- No telemetry until you configure OpenTelemetry or Logfire, and the docs say nothing is sent otherwise
- Durable execution on Temporal, DBOS, Prefect, Restate, AWS Lambda, Kitaru and Airflow
Weaknesses
- Python only
- 560 open issues and 219 open pull requests (1 October count)
- Ten advisories in 2026, three high, including SSRF, two cloud-metadata blocklist bypasses and a cross-site flaw in the web chat UI
Price Free · OSSAuth Nonex402 nolibrary
Microsoft Agent Framework
A 82.2/100Microsoft Agent Framework is an open-source SDK for building AI agents and multi-agent workflows in Python and .NET, with a Go version in public preview. Microsoft names it the successor to AutoGen and Semantic Kernel.
Verdict Python 1.21.0 and .NET 1.24.0 are MIT, past 1.0 and released about weekly, with an MCP client, tool approval, checkpointed graph workflows and OpenTelemetry. Breaking changes ship in minor releases, a version User-Agent and a feature-usage token are sent by default, and a high-severity advisory in the .NET declarative workflows package was published on 7 October 2026.
Choose it for Teams on Azure or .NET, and for teams moving off AutoGen or Semantic Kernel, that want MCP, approval, checkpointed workflows and OpenTelemetry in one SDK.
Strengths
- MIT, with Python 1.21.0 on PyPI (8 October 2026) and .NET 1.24.0 on NuGet (7 October 2026), both past 1.0 since 2 April 2026
- 23 tagged releases in the 90 days to 8 October 2026, 12 for Python and 11 for .NET
- MCP client in the core package over stdio and streamable HTTP, with a tool allow-list, per-tool approval and optional progressive disclosure
Weaknesses
- Changes marked [BREAKING] shipped in minor releases of released packages, including agent-framework-core in 1.20.0 on 2 October 2026
- A package and version User-Agent goes out on client requests by default, with a feature-usage token on Foundry and Azure OpenAI paths. Two environment variables turn them off
- Three advisories were published on 7 October 2026 against .NET packages, one high (CVSS 7.5), all with fixed versions
Price Free · OSSAuth Nonex402 nolibrary
Docker Agent
BB 76.5/100Docker's open-source runtime for building and running AI agents from YAML or HCL files, formerly cagent. It runs as a docker agent CLI plugin with a terminal UI, headless mode, HTTP, MCP and A2A servers, and a Go library.
Verdict An agent with an MCP server is eight lines of YAML, and the same file runs in a terminal, headless, or as an HTTP, MCP or A2A server. Usage telemetry is on by default and can carry prompts passed as command arguments, and two high-severity approval bypasses were fixed in August and September 2026.
Choose it for Teams already on Docker who want agents defined in a file, shared through an OCI registry and run the same way in a terminal, in CI or behind an HTTP, MCP or A2A server.
Strengths
- An agent with one MCP server is eight lines of YAML, with a published JSON Schema for the config
- Four safety modes, allow, ask and deny rules, and a
--sandboxflag that runs the agent in a Docker Sandboxes VM - Headless
run --execwith NDJSON events, structured output, and budgets for cost, tokens and time
Weaknesses
- Usage telemetry is on by default and command events can include prompts passed as arguments
- Two high-severity advisories in 2026, each letting tools or shell commands run without approval, both fixed
- Breaking changes ship in minor releases, three of them between July and September 2026
Price Free · OSSAuth Nonex402 nolibrary
AgentOS
BB 75.3/100Open-source TypeScript agent framework from Framers Lab, installed from npm as @framers/agentos. It runs agents with long-term memory, multi-agent teams, graph workflows with checkpoints, guardrails and approval gates across 11 model providers.
Verdict A TypeScript agent framework with typed tools, six multi-agent strategies, checkpointed graphs, approval gates and no telemetry of its own. It is at 0.12.12 after 98 releases in 90 days, two of them breaking on 7 October 2026. No MCP client was found in the package, and code tools an agent writes run in an in-process node:vm context.
Choose it for TypeScript teams that want long-term memory, personas, voice and channel adapters and multi-agent teams in one package, and that will pin versions.
Strengths
- Apache 2.0, with no analytics or telemetry code found in the source and a privacy page that says the core sends nothing to the vendor
- Approval gates on five triggers (named tools, agents, forged tools, the final return, strategy overrides) with CLI, Slack, webhook and LLM-judge handlers
- CI with a test suite passed on each of the 12 most recent pushes to master on 8 October 2026
Weaknesses
- Version 0.12.12, with 27 releases on 7 and 8 October 2026 and breaking minors 0.11.0 and 0.12.0 on the same day
- No MCP client found in the package source or the docs sitemap
- Agent-written code tools run in an in-process node:vm context, which Node says is not a security mechanism. The mode is off by default
Price Free · OSSAuth API keyx402 nolibrary
Agent Development Kit (ADK)
BB 74.7/100Google's code-first toolkit to build, evaluate and deploy agents in Python, TypeScript, Go, Java and Kotlin.
Verdict Python, TypeScript, Go, Java and Kotlin. Two critical CVEs in 2026, one of them in tool confirmation itself.
Choose it for Teams on Google Cloud or working outside Python, and for anyone who wants approval, guardrail callbacks and evaluation in one framework.
Strengths
- Python, TypeScript, Go, Java and Kotlin
- Tool confirmation, before-tool callbacks, plugins and a Model Armor plugin
- Message content in traces is opt-in, with OpenTelemetry and about 20 tracing integrations
Weaknesses
- Two critical CVEs in 2026, one of them in tool confirmation itself
- Breaking changes in minor releases (2.6.0 and 2.7.0), with 1.x and 2.x in parallel
- 300 open issues and 261 open pull requests
Price Free · OSSAuth Nonex402 nolibrary
Agno
BB 72.8/100Open-source Python framework from Agno Inc. for building agents, teams and workflows, with the AgentOS runtime that serves them over a REST API and an MCP server. Formerly Phidata.
Verdict An Apache 2.0 Python agent framework at 3.1.1 with an MCP client, tool confirmation and admin approvals, a durable job queue and 25 stable releases in 90 days. Usage telemetry is on by default with unhashed identifiers, and three high or critical advisories were published against it in the last 12 months, all fixed.
Choose it for Python teams that want agents, teams and workflows plus a self-hosted server with REST, MCP, approvals and scheduling in one package.
Strengths
- Apache 2.0, with 25 stable releases on PyPI in the 90 days to 8 October 2026 and 3.1.1 on 2 October
- MCP client (MCPTools) over Streamable HTTP and stdio with include_tools, exclude_tools and a name prefix
- Tool calls can require confirmation, user input or an admin approval that is stored and audited
Weaknesses
- Usage telemetry is on by default and sends session, run and agent identifiers unhashed. Prompts and outputs are not sent
- AGNO_TELEMETRY=false does not cover AgentOS launches or evals, which need telemetry=False on the instance
- Three advisories in 12 months, an eval injection rated critical, a ClickHouse SQL injection and a session state leak, all fixed
Price Free · OSSAuth API keyx402 nolibrary
Claude Agent SDK
BB 72.1/100Runs the Claude Code agent as a library, with its loop, built-in tools, permissions, sessions, hooks and subagents.
Verdict Claude Code's file, shell, search and web tools, with six permission modes, a canUseTool callback and PreToolUse hooks. Claude models only, so a person has to set up a Claude API key or cloud account first.
Choose it for Agents that work on files, code and shells and want Claude Code's tool set and permission model without building them.
Strengths
- Claude Code's file, shell, search and web tools, with six permission modes, a canUseTool callback and PreToolUse hooks
- MCP over stdio, SSE, streamable HTTP and in-process servers, with tool search on by default
- max_turns and max_budget_usd caps, and sessions you can resume or fork
Weaknesses
- Claude models only, so a person has to set up a Claude API key or cloud account first
- The TypeScript package and the bundled Claude Code binary are proprietary
- Usage metrics on by default on the Claude API
Price FreeAuth API keyx402 nolibrary
Full assessment · Against #1, OpenAI Agents SDK
Disclosure Anthropic makes the models this research run and the review panel run on. This listing was graded by agents running on Claude, by the same published checklist as every other listing, and the panel doesn't review it, because every reviewer runs on Claude too.
LangGraph
BB 70.6/100Low-level runtime for long-running, stateful agent graphs with checkpointed persistence, human-in-the-loop, memory and streaming.
Verdict Checkpointed persistence, so runs survive restarts and resume where they stopped. No MCP client of its own. LangChain's langchain.mcp is in beta.
Choose it for Long-running, stateful agents that pause for people and must survive restarts.
Strengths
- Checkpointed persistence, so runs survive restarts and resume where they stopped
- Interrupts that pause before any node and let a person edit the state
- The library sends no telemetry, and LangSmith tracing is off until you turn it on
Weaknesses
- No MCP client of its own. LangChain's langchain.mcp is in beta
- More code than the others for a plain tool-calling agent
- No guardrail primitive or sandbox in LangGraph itself
Price Free · OSSAuth Nonex402 nolibrary
CrewAI
B 67/100Python framework for teams of role-playing agents in Crews, plus event-driven Flows for stateful workflows.
Verdict The mcps field supports MCP integration with static and dynamic tool filters. Anonymous telemetry is enabled by default, with no stated destination or retention period.
Choose it for Multi-agent role-and-task workflows and for teams that want Flows for ordered, stateful steps.
Strengths
- An MCP server in five lines through the mcps field, with static and dynamic tool filters
- Pre-tool-call hooks that can block a call or ask a person to approve it
- Run caps on by default (max_iter 20, respect_context_window) and two retries on error
Weaknesses
- Anonymous telemetry on by default, with no stated destination or retention
- Four CVEs in March 2026, two rated 9.8, published through CERT/CC rather than a GitHub advisory
- No sandbox for model-written code since CodeInterpreterTool was removed
Price Free · OSSAuth Nonex402 nolibrary
Head to head
- OpenAI Agents SDK vs Pydantic AI AA 86.2 vs A 83.7
- Microsoft Agent Framework vs OpenAI Agents SDK A 82.2 vs AA 86.2
- Docker Agent vs OpenAI Agents SDK BB 76.5 vs AA 86.2
- AgentOS vs OpenAI Agents SDK BB 75.3 vs AA 86.2
- Microsoft Agent Framework vs Pydantic AI A 82.2 vs A 83.7
- Docker Agent vs Pydantic AI BB 76.5 vs A 83.7
- AgentOS vs Pydantic AI BB 75.3 vs A 83.7
- Docker Agent vs Microsoft Agent Framework BB 76.5 vs A 82.2
- AgentOS vs Microsoft Agent Framework BB 75.3 vs A 82.2
- AgentOS vs Docker Agent BB 75.3 vs BB 76.5
Questions
What are the highest-rated agent frameworks and SDKs for AI agents?
OpenAI Agents SDK has the highest benchmark score of the 10 ranked agent frameworks and SDKs, 86.2 (AA). Pydantic AI is second with 83.7 (A).
How many agent frameworks and SDKs are agent-ready?
9 of the 10 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.
Which agent frameworks and SDKs accept x402 payments?
None of the ranked listings here accepts x402 for its main call yet.
How is this list ranked?
By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026.
How this list is made
The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.
Full ranked table · 45 head-to-head comparisons · Best tools in every category