{
  "fixes": {
    "slug": "skyvern",
    "name": "Skyvern",
    "listing": "https://www.anchorterminal.com/tools/skyvern",
    "markdown": "# Fix list: Skyvern\n\nFrom Anchor Terminal's listing at https://www.anchorterminal.com/tools/skyvern, the October 2026 research run, assessed 8 October 2026. Grade B, 63.1 out of 100.\n\nThis is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.\n\nFor a coding agent working on Skyvern: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.\n\n## 1. Reliability, 57 out of 100, up to 8.6 more on the total\n\nWhy it scored 57: Graded as a hosted service, Skyvern Cloud at api.skyvern.com. status.skyvern.com is a Statuspage site with three components (API, web application, async workers), 90-day uptime bars and an incident feed back to March 2025 (20). The feed has one incident in the 90 days to 8 October 2026, marked major by Skyvern, on 6 August, when an upstream model provider rejected requests and task and workflow runs failed at elevated rates from about 21:20 UTC to 22:30, roughly 70 minutes. A 29-minute run of 503s on 3 July falls just outside the window (10). The pricing page gives concurrent runs per plan (1, 10, 25, 100) and API responses carry `ratelimit-policy: \"submit-run\";q=50;w=60`, but we found no documented request limits (8). The OpenAPI document describes 503 with `Retry-After` on run submission and 429 on two recipe endpoints, and the SDKs retry network errors and 5xx with backoff. `Idempotency-Key` exists only on POST /v1/agents, so retried run submissions rely on the 503 saying no run was created (9). The pricing FAQ mentions custom SLAs for Enterprise and nothing is published (0). `/v1` is declared stable (10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):\n\nHosted APIs, MCP servers, models and platforms.\n\n- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.\n\nLocal packages, SDKs, frameworks and stdio MCP servers.\n\n- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.\n\nProtocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.\n\n## 2. Payments \u0026 pricing, 32 out of 100, up to 8.5 more on the total\n\nWhy it scored 32: No x402, MPP or L402 in the docs, llms.txt or pricing page (0). Plan prices are public with credit allowances, Free 5,000 credits once, Hobby $29 for 30,000 a month, Pro $149 for 150,000, Enterprise custom. The price of extra credits isn't published, and the pages disagree on what a credit buys, one credit an action in the billing docs against about 170 or about 200 actions for 5,000 credits (12). The Free plan needs no card, per the pricing page (20). Signup is a browser flow, including `skyvern signup` from the CLI, so a person has to create the account (0). The open-source server is free to run under AGPL-3.0 with your own model keys, and we scored the hosted service because that is what the MCP and API docs point agents to.\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):\n\nThe published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).\n\n- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).\n\nPayment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.\n\nOpen-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.\n\n## 3. Security \u0026 auth, 59 out of 100, up to 7.2 more on the total\n\nWhy it scored 59: An organisation-wide API key in the `x-api-key` header, revocable and rotated from Settings, or OAuth 2.0 with PKCE, dynamic client registration and rotating single-use refresh tokens. The docs say plainly that OAuth scopes are identity claims and that there are no read-only, per-endpoint or per-resource credentials, so we scored plain revocable keys. No secret travels in a URL (20). MCP tool scopes are documented as a usability filter and not a permission boundary. Separate organisations are the only isolation, Human Interaction blocks are listed for Enterprise, and `skyvern_act` rejects prompts that contain passwords (6). Web pages are untrusted content. Stored credentials reach the browser directly and appear to the model as placeholders, the May 2026 changelog records sanitising of page content in prompt templates, and one tool description says page content is data, not instructions. We found no guidance page on prompt injection (8). GET /v1/audit-events/export returns organisation audit events for a default 90 days as JSON or CSV, and each run keeps a recording, screenshots, an action timeline and a HAR file (14). The site claims SOC 2 Type II and HIPAA on Enterprise, with a trust centre that didn't render for us. SECURITY.md routes reports to GitHub private advisories and its supported-versions table still says 0.1.x. No security.txt and no bug bounty found. The webhook page warns that its verifier examples before August 2026 accepted any signature of the right length (11).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-security):\n\n- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.\n\nModels are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.\n\n## 4. Agent ergonomics, 76 out of 100, up to 3.9 more on the total\n\nWhy it scored 76: The full MCP surface is 114 tools by the docs' count (5). Scopes cut that to 29 (`operate`), 32 (`lean`), 54 (`browser`) or 60 (`build`) through a URL such as /mcp/x/lean or the `X-Skyvern-Scope` header, and page reads are size-capped and selector-scoped (9, so 14). List endpoints take `page` and `page_size`, with `status`, `search_key`, tag and date filters, and `max_steps` bounds a run (17). MCP results return codes such as SELECTOR_NOT_FOUND and SESSION_EXPIRED with a hint, runs report `status`, `failure_reason` and a caller-defined `error_code_mapping`, and the SDKs raise typed errors. HTTP error bodies are mostly unspecified (16). Every tool has readOnlyHint, destructiveHint and openWorldHint annotations, 117 registrations in the 1.0.55 source. `Idempotency-Key` covers agent creation only, not run submission (14). A task needs only a prompt, and official SDKs exist for Python and TypeScript (15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):\n\n- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.\n\nModels are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.\n\n## 5. Transparency \u0026 trust, 72 out of 100, up to 2.5 more on the total\n\nMade of editorial 66, provenance 78.\n\nWhy it scored 72: The core is AGPL-3.0, an OSI licence, in a public repository. The README says anti-bot measures are in the managed cloud only (27). The privacy policy (last modified 1 October 2026) keeps data \"as long as necessary\" with deletion on request and no stated periods. The terms (27 February 2026) say data may go to third-party model providers such as OpenAI and Anthropic and that anonymised data may train models unless you opt out by email. Section 7.3 of the terms says the services aren't designed to comply with HIPAA, while the pricing page lists HIPAA compliance on Enterprise. No public DPA found (12). The deprecation policy promises at least six months' notice, 12 for a major version, a changelog entry, `Deprecation` and `Sunset` headers and `deprecated` flags in the spec (20). Model providers are named only as examples, artifact URLs point to Amazon S3, and no subprocessor list or data location statement was readable. The open-source server discloses PostHog usage telemetry in its README with `SKYVERN_TELEMETRY=false` to turn it off (7).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):\n\n- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).\n\nThe other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.\n\nProvenance checks not met in full (half of this category, computed from checked facts):\n\n- Domain age: skyvern.com, registered 2023-10-16 (2 years) (7 of 15)\n- Terms of service: read, states 6 of the 7 things a reader expects (9.1 of 10)\n- Privacy policy: read, states 7 of the 8 things a reader expects, and has 1 clause that costs points (7.3 of 10)\n- security.txt: not found (0 of 10)\n\n## 6. Schema \u0026 documentation, 88 out of 100, up to 2 more on the total\n\nWhy it scored 88: OpenAPI 3.1.0 at /docs/api-reference/openapi.json with 94 operations on 71 paths and 261 schemas, and MCP tools with typed parameters (25). llms.txt, llms-full.txt and a Markdown twin of every docs page (10). The MCP server ships a routing table that says which tool to use for each kind of job, what each costs in model calls and what not to use it for, and tool docstrings repeat it. Many REST operations have one-line descriptions such as \"Run a task\" (16). Inputs are typed with enums and ranges, such as `limit` 1 to 500 on the audit export, but `skyvern_workflow_create` takes the whole definition as a JSON or YAML string and the spec marks `x-api-key` as optional on every operation (11). Code samples in Python, TypeScript and cURL and a full error-handling guide, while 422 is the only error documented on most operations, with 404 on 31 of 94 and 429 on two (11). `/v1` path versioning with a written compatibility policy and a weekly dated changelog (15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):\n\nAPIs and MCP servers.\n\n- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.\n\nModels are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.\n\n## 7. Maintenance \u0026 community, 85 out of 100, up to 1.3 more on the total\n\nWhy it scored 85: Version 1.0.55 on PyPI and npm on 1 October 2026, seven days before this check (30). Twelve weekly changelog entries since 13 July 2026, and PyPI releases 1.0.47, 1.0.48 and 1.0.55 in the same period (20). The repository had 200 commits between 16 September and 7 October and 48 open issues, several opened in the last fortnight. GitHub didn't show us the replies, and an issue of 8 September 2026 asks whether anyone watches the private advisory queue, so we scored this below full (14). Listed in the official MCP registry as io.github.Skyvern-AI/skyvern, active, but the entry is version 1.0.23 from 13 March 2026 and lists only the API-key header (13). CI runs pre-commit hooks, a migration check, pytest and pip smoke tests on Python 3.11 and 3.13, with a locked dependency file. We couldn't read whether the default branch passes (8).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):\n\n- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.\n\nModels are read for deprecation notice periods and model churn rather than release counts.\n\n## Deductions\n\nEach comes off the total. A fixed and documented problem counts for less at the next check.\n\n- 2026-06-30. A public issue reports non-blind SSRF from workflow `http_request` and file download blocks to loopback, private and metadata addresses in version 1.0.39, with a reproduction (https://github.com/Skyvern-AI/skyvern/issues/6915). The issue is still open on 8 October 2026 and no advisory is published, but the 1.0.55 source routes these requests through an SSRF-guarded resolver, so we deduct 3 and not more. We didn't reproduce it, and whether Skyvern Cloud was exposed isn't stated.\n\n## What we couldn't check\n\nWhat we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.\n\n- unchecked: the trust centre at trust.skyvern.com is a JavaScript application that didn't render for our reader, so the SOC 2 Type II report, the HIPAA statement and any subprocessor list rest on the pricing and home pages\n- unchecked: reply times on GitHub issues and whether CI passes on the default branch. The GitHub API refused us for rate limits and the issue pages didn't show comments\n- unchecked: whether the SSRF reported in issue 6915 affected Skyvern Cloud, and which release added the SSRF-guarded resolver. The changelog doesn't mention it\n- What a credit buys. The billing docs say one credit a browser action, the pricing page says 5,000 credits is about 170 actions, the cost-control page about 200, and the pricing FAQ says credits depend on run complexity and duration\n- Whether the `ratelimit-policy` header (50 run submissions in 60 seconds) is the enforced limit on every plan. It isn't documented\n- Enterprise concurrency is 100 on the pricing page and unlimited in the billing docs\n- Whether the permissive webhook verifier examples before August 2026 merit a deduction. We recorded them in the security note and didn't deduct, because the page discloses the fault and we couldn't date the fix\n- The audit export says only organisation admins and full API keys may export, which implies a restricted key type that the authentication page says doesn't exist\n- firstReleased is left empty. PyPI's earliest file is 0.1.53 from 6 February 2025, and the repository is older\n\n## Weaknesses\n\n- Every API key and OAuth token has full organisation authority. The docs state there are no read-only, per-endpoint or per-resource keys\n- No request rate limits in the docs. Responses carry `ratelimit-policy: \"submit-run\";q=50;w=60`, and only plan concurrency is published\n- What a credit buys is stated three ways (one credit an action, 5,000 credits for about 170 or about 200 actions, by run complexity)\n- An SSRF report against 1.0.39 from 30 June 2026 is still open with no advisory, though 1.0.55 source has SSRF guards\n- Retention is \"as long as necessary\", anonymised data may train models unless you opt out by email, and no subprocessor list was readable\n\n## What costs an agent a turn today\n\nThe notes we give agents before they call it. Each one is a workaround an agent shouldn't need.\n\n- Connect to https://api.skyvern.com/mcp/x/lean or /x/operate, or send `X-Skyvern-Scope`, so the client loads 32 or 29 tools. The scope filters the list and does not limit permissions\n- Use a separate Skyvern organisation for each blast radius. Any key or OAuth token can read and write stored credentials and delete workflows\n- Never pass passwords to `skyvern_act` or `skyvern_type`. Store them as credentials and call `skyvern_login`\n- Set `max_steps` on tasks, or `x-max-steps-override` on agent runs, to cap credits. A run that reaches the cap ends as `timed_out`\n- On 503 from POST /v1/run/agents, wait `Retry-After` seconds. No run was created, so resubmitting is safe\n\n## When it's done\n\nSend what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `\"kind\": \"dispute\"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.\n",
    "grade": "B",
    "score": 63.1,
    "assessed": "2026-10-08",
    "run": "October 2026 research run",
    "categories": [
      {
        "key": "reliability",
        "name": "Reliability",
        "score": 57,
        "maxGain": 8.6,
        "reason": "Graded as a hosted service, Skyvern Cloud at api.skyvern.com. status.skyvern.com is a Statuspage site with three components (API, web application, async workers), 90-day uptime bars and an incident feed back to March 2025 (20). The feed has one incident in the 90 days to 8 October 2026, marked major by Skyvern, on 6 August, when an upstream model provider rejected requests and task and workflow runs failed at elevated rates from about 21:20 UTC to 22:30, roughly 70 minutes. A 29-minute run of 503s on 3 July falls just outside the window (10). The pricing page gives concurrent runs per plan (1, 10, 25, 100) and API responses carry `ratelimit-policy: \"submit-run\";q=50;w=60`, but we found no documented request limits (8). The OpenAPI document describes 503 with `Retry-After` on run submission and 429 on two recipe endpoints, and the SDKs retry network errors and 5xx with backoff. `Idempotency-Key` exists only on POST /v1/agents, so retried run submissions rely on the 503 saying no run was created (9). The pricing FAQ mentions custom SLAs for Enterprise and nothing is published (0). `/v1` is declared stable (10).",
        "checklist": [
          "Hosted APIs, MCP servers, models and platforms.",
          "- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.",
          "Local packages, SDKs, frameworks and stdio MCP servers.",
          "- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.",
          "Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-reliability"
      },
      {
        "key": "payments",
        "name": "Payments \u0026 pricing",
        "score": 32,
        "maxGain": 8.5,
        "reason": "No x402, MPP or L402 in the docs, llms.txt or pricing page (0). Plan prices are public with credit allowances, Free 5,000 credits once, Hobby $29 for 30,000 a month, Pro $149 for 150,000, Enterprise custom. The price of extra credits isn't published, and the pages disagree on what a credit buys, one credit an action in the billing docs against about 170 or about 200 actions for 5,000 credits (12). The Free plan needs no card, per the pricing page (20). Signup is a browser flow, including `skyvern signup` from the CLI, so a person has to create the account (0). The open-source server is free to run under AGPL-3.0 with your own model keys, and we scored the hosted service because that is what the MCP and API docs point agents to.",
        "checklist": [
          "The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).",
          "- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).",
          "Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.",
          "Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-payments"
      },
      {
        "key": "security",
        "name": "Security \u0026 auth",
        "score": 59,
        "maxGain": 7.2,
        "reason": "An organisation-wide API key in the `x-api-key` header, revocable and rotated from Settings, or OAuth 2.0 with PKCE, dynamic client registration and rotating single-use refresh tokens. The docs say plainly that OAuth scopes are identity claims and that there are no read-only, per-endpoint or per-resource credentials, so we scored plain revocable keys. No secret travels in a URL (20). MCP tool scopes are documented as a usability filter and not a permission boundary. Separate organisations are the only isolation, Human Interaction blocks are listed for Enterprise, and `skyvern_act` rejects prompts that contain passwords (6). Web pages are untrusted content. Stored credentials reach the browser directly and appear to the model as placeholders, the May 2026 changelog records sanitising of page content in prompt templates, and one tool description says page content is data, not instructions. We found no guidance page on prompt injection (8). GET /v1/audit-events/export returns organisation audit events for a default 90 days as JSON or CSV, and each run keeps a recording, screenshots, an action timeline and a HAR file (14). The site claims SOC 2 Type II and HIPAA on Enterprise, with a trust centre that didn't render for us. SECURITY.md routes reports to GitHub private advisories and its supported-versions table still says 0.1.x. No security.txt and no bug bounty found. The webhook page warns that its verifier examples before August 2026 accepted any signature of the right length (11).",
        "checklist": [
          "- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.",
          "Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-security"
      },
      {
        "key": "ergonomics",
        "name": "Agent ergonomics",
        "score": 76,
        "maxGain": 3.9,
        "reason": "The full MCP surface is 114 tools by the docs' count (5). Scopes cut that to 29 (`operate`), 32 (`lean`), 54 (`browser`) or 60 (`build`) through a URL such as /mcp/x/lean or the `X-Skyvern-Scope` header, and page reads are size-capped and selector-scoped (9, so 14). List endpoints take `page` and `page_size`, with `status`, `search_key`, tag and date filters, and `max_steps` bounds a run (17). MCP results return codes such as SELECTOR_NOT_FOUND and SESSION_EXPIRED with a hint, runs report `status`, `failure_reason` and a caller-defined `error_code_mapping`, and the SDKs raise typed errors. HTTP error bodies are mostly unspecified (16). Every tool has readOnlyHint, destructiveHint and openWorldHint annotations, 117 registrations in the 1.0.55 source. `Idempotency-Key` covers agent creation only, not run submission (14). A task needs only a prompt, and official SDKs exist for Python and TypeScript (15).",
        "checklist": [
          "- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.",
          "Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-ergonomics"
      },
      {
        "key": "transparency",
        "name": "Transparency \u0026 trust",
        "score": 72,
        "maxGain": 2.5,
        "reason": "The core is AGPL-3.0, an OSI licence, in a public repository. The README says anti-bot measures are in the managed cloud only (27). The privacy policy (last modified 1 October 2026) keeps data \"as long as necessary\" with deletion on request and no stated periods. The terms (27 February 2026) say data may go to third-party model providers such as OpenAI and Anthropic and that anonymised data may train models unless you opt out by email. Section 7.3 of the terms says the services aren't designed to comply with HIPAA, while the pricing page lists HIPAA compliance on Enterprise. No public DPA found (12). The deprecation policy promises at least six months' notice, 12 for a major version, a changelog entry, `Deprecation` and `Sunset` headers and `deprecated` flags in the spec (20). Model providers are named only as examples, artifact URLs point to Amazon S3, and no subprocessor list or data location statement was readable. The open-source server discloses PostHog usage telemetry in its README with `SKYVERN_TELEMETRY=false` to turn it off (7).",
        "blend": "editorial 66, provenance 78",
        "checklist": [
          "- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).",
          "The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-transparency"
      },
      {
        "key": "schema",
        "name": "Schema \u0026 documentation",
        "score": 88,
        "maxGain": 2,
        "reason": "OpenAPI 3.1.0 at /docs/api-reference/openapi.json with 94 operations on 71 paths and 261 schemas, and MCP tools with typed parameters (25). llms.txt, llms-full.txt and a Markdown twin of every docs page (10). The MCP server ships a routing table that says which tool to use for each kind of job, what each costs in model calls and what not to use it for, and tool docstrings repeat it. Many REST operations have one-line descriptions such as \"Run a task\" (16). Inputs are typed with enums and ranges, such as `limit` 1 to 500 on the audit export, but `skyvern_workflow_create` takes the whole definition as a JSON or YAML string and the spec marks `x-api-key` as optional on every operation (11). Code samples in Python, TypeScript and cURL and a full error-handling guide, while 422 is the only error documented on most operations, with 404 on 31 of 94 and 429 on two (11). `/v1` path versioning with a written compatibility policy and a weekly dated changelog (15).",
        "checklist": [
          "APIs and MCP servers.",
          "- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.",
          "Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-schema"
      },
      {
        "key": "maintenance",
        "name": "Maintenance \u0026 community",
        "score": 85,
        "maxGain": 1.3,
        "reason": "Version 1.0.55 on PyPI and npm on 1 October 2026, seven days before this check (30). Twelve weekly changelog entries since 13 July 2026, and PyPI releases 1.0.47, 1.0.48 and 1.0.55 in the same period (20). The repository had 200 commits between 16 September and 7 October and 48 open issues, several opened in the last fortnight. GitHub didn't show us the replies, and an issue of 8 September 2026 asks whether anyone watches the private advisory queue, so we scored this below full (14). Listed in the official MCP registry as io.github.Skyvern-AI/skyvern, active, but the entry is version 1.0.23 from 13 March 2026 and lists only the API-key header (13). CI runs pre-commit hooks, a migration check, pytest and pip smoke tests on Python 3.11 and 3.13, with a locked dependency file. We couldn't read whether the default branch passes (8).",
        "checklist": [
          "- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.",
          "Models are read for deprecation notice periods and model churn rather than release counts."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-maintenance"
      }
    ],
    "provenance": [
      {
        "label": "Domain age",
        "value": "skyvern.com, registered 2023-10-16 (2 years)",
        "points": 7,
        "max": 15
      },
      {
        "label": "Terms of service",
        "value": "read, states 6 of the 7 things a reader expects",
        "points": 9.1,
        "max": 10
      },
      {
        "label": "Privacy policy",
        "value": "read, states 7 of the 8 things a reader expects, and has 1 clause that costs points",
        "points": 7.3,
        "max": 10
      },
      {
        "label": "security.txt",
        "value": "not found",
        "points": 0,
        "max": 10
      }
    ],
    "deductions": [
      "2026-06-30. A public issue reports non-blind SSRF from workflow `http_request` and file download blocks to loopback, private and metadata addresses in version 1.0.39, with a reproduction (https://github.com/Skyvern-AI/skyvern/issues/6915). The issue is still open on 8 October 2026 and no advisory is published, but the 1.0.55 source routes these requests through an SSRF-guarded resolver, so we deduct 3 and not more. We didn't reproduce it, and whether Skyvern Cloud was exposed isn't stated."
    ],
    "unchecked": [
      "unchecked: the trust centre at trust.skyvern.com is a JavaScript application that didn't render for our reader, so the SOC 2 Type II report, the HIPAA statement and any subprocessor list rest on the pricing and home pages",
      "unchecked: reply times on GitHub issues and whether CI passes on the default branch. The GitHub API refused us for rate limits and the issue pages didn't show comments",
      "unchecked: whether the SSRF reported in issue 6915 affected Skyvern Cloud, and which release added the SSRF-guarded resolver. The changelog doesn't mention it",
      "What a credit buys. The billing docs say one credit a browser action, the pricing page says 5,000 credits is about 170 actions, the cost-control page about 200, and the pricing FAQ says credits depend on run complexity and duration",
      "Whether the `ratelimit-policy` header (50 run submissions in 60 seconds) is the enforced limit on every plan. It isn't documented",
      "Enterprise concurrency is 100 on the pricing page and unlimited in the billing docs",
      "Whether the permissive webhook verifier examples before August 2026 merit a deduction. We recorded them in the security note and didn't deduct, because the page discloses the fault and we couldn't date the fix",
      "The audit export says only organisation admins and full API keys may export, which implies a restricted key type that the authentication page says doesn't exist",
      "firstReleased is left empty. PyPI's earliest file is 0.1.53 from 6 February 2025, and the repository is older"
    ],
    "weaknesses": [
      "Every API key and OAuth token has full organisation authority. The docs state there are no read-only, per-endpoint or per-resource keys",
      "No request rate limits in the docs. Responses carry `ratelimit-policy: \"submit-run\";q=50;w=60`, and only plan concurrency is published",
      "What a credit buys is stated three ways (one credit an action, 5,000 credits for about 170 or about 200 actions, by run complexity)",
      "An SSRF report against 1.0.39 from 30 June 2026 is still open with no advisory, though 1.0.55 source has SSRF guards",
      "Retention is \"as long as necessary\", anonymised data may train models unless you opt out by email, and no subprocessor list was readable"
    ],
    "agentNotes": [
      "Connect to https://api.skyvern.com/mcp/x/lean or /x/operate, or send `X-Skyvern-Scope`, so the client loads 32 or 29 tools. The scope filters the list and does not limit permissions",
      "Use a separate Skyvern organisation for each blast radius. Any key or OAuth token can read and write stored credentials and delete workflows",
      "Never pass passwords to `skyvern_act` or `skyvern_type`. Store them as credentials and call `skyvern_login`",
      "Set `max_steps` on tasks, or `x-max-steps-override` on agent runs, to cap credits. A run that reaches the cap ends as `timed_out`",
      "On 503 from POST /v1/run/agents, wait `Retry-After` seconds. No run was created, so resubmitting is safe"
    ],
    "recheck": "https://www.anchorterminal.com/builders/#disputes"
  },
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-08",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  }
}
