Category · Agent frameworks

Agent frameworks and SDKs

Libraries that run the agent loop. Tool calling, MCP, multi-agent hand-offs, durable state, human approval and tracing. Compared on what they do by default, including the telemetry they send.

Capability keys agent.framework · agent.multi-agent · agent.durable · agent.mcp-client · All tools

letme.dev/agent.framework picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

6listings graded
5agent-ready (BB+)
28desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Auth
Pricing
Status
6 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
1 OpenAI Agents SDKOpenAI · Agent framework Frameworks AA 86.5 3.9 (8) Free · OSS
7 Pydantic AIPydantic · Agent framework Frameworks A 80 4.0 (8) Free · OSS
45 Agent Development Kit (ADK)Google · Agent framework Frameworks BB 74.9 3.0 (8) Free · OSS
71 Claude Agent SDKAnthropic · Agent framework Frameworks BB 72.4 none Free
95 LangGraphLangChain · Agent framework Frameworks BB 70.6 3.5 (2) Free · OSS
149 CrewAICrewAI · Agent framework Frameworks B 67 3.0 (2) Free · OSS

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

Indexed, not reviewed (2)

Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.

ListingKindWhat it doesWhy it's here
FlurryPORT
flurryport.io
MCP serverWebhook capture and replay, delivery pipes, and signed multi-agent rooms. No signup to start.vendor's own
Gemot Deliberation Server
gemot.dev
MCP serverDeliberation primitive for multi-agent coordination: cruxes, vote clustering, consensus.vendor's own

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.