{
  "fixes": {
    "slug": "docker-model-runner",
    "name": "Docker Model Runner",
    "listing": "https://www.anchorterminal.com/tools/docker-model-runner",
    "markdown": "# Fix list: Docker Model Runner\n\nFrom Anchor Terminal's listing at https://www.anchorterminal.com/tools/docker-model-runner, the October 2026 research run, assessed 8 October 2026. Grade C, 57.1 out of 100.\n\nThis is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public.\n\nFor a coding agent working on Docker Model Runner: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published.\n\n## 1. Security \u0026 auth, 40 out of 100, up to 10.5 more on the total\n\nWhy it scored 40: Read with the tool checklist, for the local API. No credential, by design. The docs say the API is not authenticated and that any client that can reach it, including other containers on the same Docker network, can pull, load and run models. Host-side TCP is off by default in Docker Desktop and on by default on port 12434 in Docker Engine, and cross-origin requests are refused with 403 unless the origin is localhost, 127.0.0.1, 0.0.0.0 or listed in `DMR_ORIGINS` (7 of 30). No read-only mode or per-caller limit. Runtime flags pass an allowlist, engines run sandboxed on macOS and Windows and in a container on Linux, and Enhanced Container Isolation, a Docker Business control, blocks container access (6 of 20). The API returns model output, with no injection guidance in the docs (6 of 15). `docker model requests` and the Requests tab show recent requests and responses, the source keeps the last 10 per model, and `/metrics` exposes Prometheus counters. No caller identity (8 of 15). www.docker.com has a valid security.txt with a disclosure policy, SECURITY.md promises an acknowledgement within 72 hours, and two advisories were published with CVEs in 2026. No monetary bounty for this project (13 of 20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-security):\n\n- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.\n\nModels are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing.\n\n## 2. Schema \u0026 documentation, 49 out of 100, up to 8.3 more on the total\n\nWhy it scored 49: Read for an API. No OpenAPI or similar file was found in the repository or the docs. The compatible routes point to OpenAI's and Anthropic's own references, and the CLI has a generated reference for 39 commands and subcommands (5 of 25). docs.docker.com has llms.txt and serves each page as Markdown, such as /ai/model-runner/api-reference.md (10). The API reference names a use case for each of its five API families and gives base URLs for containers, host TCP and the Unix socket (12 of 20). Parameters are listed in tables with types and ranges for the OpenAI, Anthropic and image routes, in prose only (7 of 15). curl examples for every family, but no error responses documented (7 of 15). Dated GitHub releases with notes. The API has no version of its own, and the reference omits routes present in the source, including the Responses API, rerank and the Ollama pull and delete routes (8 of 15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-schema):\n\nAPIs and MCP servers.\n\n- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.\n\nModels are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference.\n\n## 3. Agent ergonomics, 58 out of 100, up to 6.8 more on the total\n\nWhy it scored 58: Read for an API. Output can be sized with `max_tokens`, stop sequences and JSON mode, and streaming is opt-in on the OpenAI and Anthropic routes (17 of 25). Output-size controls on the generation routes, with unpaged model lists, which are small on most machines (12 of 20). Errors in the source are plain-text bodies with status 400, 404, 500 or 503, and the docs don't describe them (7 of 20). Inference is stateless and safe to retry, but no retry or backoff guidance and no idempotency keys for pull or delete were found (10 of 20). `model` is the only required field beyond the messages or prompt. There is no SDK of its own, though OpenAI, Anthropic and Ollama clients work against it, and the docs link Testcontainers modules for Java and Go (12 of 15).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics):\n\n- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.\n\nModels are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs.\n\n## 4. Payments \u0026 pricing, 60 out of 100, up to 5 more on the total\n\nWhy it scored 60: Read with the self-hosted rule, since the API an agent calls is free software on the owner's machine. No x402, MPP or L402 in the docs or the source (0). Model Runner is Apache-2.0 with nothing to buy and no account needed for the Docker Engine plugin or the `dmr` binary, so 20, 20 and 20 on the last three lines. One qualification applies to the Docker Desktop route. Docker's subscription agreement limits free Desktop use to non-commercial open-source projects and businesses with fewer than 250 employees and under US $10,000,000 in annual revenue, and paid plans run from $11 to $24 per user a month.\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-payments):\n\nThe published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).\n\n- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).\n\nPayment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.\n\nOpen-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol.\n\n## 5. Maintenance \u0026 community, 55 out of 100, up to 3.9 more on the total\n\nWhy it scored 55: v1.2.8 was released on 12 August 2026, 57 days before the check (20 of 30). Two releases in the 90 days to 8 October 2026, v1.2.7 on 11 August and v1.2.8, against 30 or so between March and June (0 of 20). 45 commits on main since 10 July, the newest on 8 October, and each of the 20 newest open issues had at least one comment, with 41 issues and 25 pull requests open (17 of 25). No SDK of its own. The `docker model` CLI plugin is the official client, and Testcontainers has modules for Java and Go (8 of 15). Dependabot runs weekly on Go modules and GitHub Actions, actions are pinned by commit, a script bumps llama.cpp, and CI passes on main (10).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance):\n\n- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.\n\nModels are read for deprecation notice periods and model churn rather than release counts.\n\n## 6. Reliability, 85 out of 100, up to 3 more on the total\n\nWhy it scored 85: Read with the local-software lines, since the API runs on the owner's machine. Bundled with Docker Desktop, installed as `docker-model-plugin` from Docker's apt and dnf repositories, and shipped as a standalone `dmr` binary through Homebrew and winget, with platform, GPU and driver requirements stated in the docs (20). The CI workflow runs lint, race-detector tests and builds on pushes and pull requests to main, with separate end-to-end, integration and daily check workflows, and the ten most recent CI runs on main passed on 8 October 2026 (25). 41 open issues and 25 open pull requests. Each of the 20 newest open issues had at least one comment, but several are regressions or failures still open, including HTTP 500 on sequential tool calls (#1063), a Windows GPU regression after Docker Desktop 4.82.0 (#1054) and a context-size setting applied nondeterministically (#1025) (17 of 25). Semver tags with notes on each GitHub release, but no changelog file and no breaking-change section in the notes we read (8 of 15). Version 1.2.8 (15). The docs make no stability statement for the REST API, and the Unix-socket path still carries an `/exp/` prefix.\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability):\n\nHosted APIs, MCP servers, models and platforms.\n\n- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.\n\nLocal packages, SDKs, frameworks and stdio MCP servers.\n\n- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.\n\nProtocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors.\n\n## 7. Transparency \u0026 trust, 73 out of 100, up to 2.4 more on the total\n\nMade of editorial 62, provenance 84.\n\nWhy it scored 73: The editorial half. Apache-2.0 for the server, the CLI and the `dmr` binary in a public repository. Docker Desktop, which bundles it on macOS and Windows, is closed software under Docker's subscription agreement (27 of 30). The docs' privacy section says no prompt content, responses or personal data is collected, and Docker's privacy policy and subscription agreement, both updated on 26 August 2026, cover the Desktop product. No retention period specific to Model Runner was found (18 of 30). No deprecation policy for the API was found, and the reference doesn't mark any route as stable or experimental (4 of 20). Telemetry is disclosed with a link to the source. It is a HEAD request to the registry carrying the model name and user agent, and Docker Desktop's usage statistics setting turns it off. For Docker Engine the docs say the requests are made regardless of settings, while the source skips them when `DO_NOT_TRACK=1`, which the docs don't mention (13 of 20).\n\nThe checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency):\n\n- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).\n\nThe other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two.\n\nProvenance checks not met in full (half of this category, computed from checked facts):\n\n- Terms of service: read, states 7 of the 7 things a reader expects, and has 2 clauses that cost points (6 of 10)\n- Status page: not found (0 of 10)\n\n## Deductions\n\nEach comes off the total. A fixed and documented problem counts for less at the next check.\n\n- 2026-02-27. GHSA-m456-c56c-hh5c (CVE-2026-28400, 7.5). The unauthenticated `/engines/_configure` route accepted arbitrary runtime flags, so a caller, including a container on Docker Desktop, could overwrite files the runner could reach, the Desktop VM disk among them. Fixed in Model Runner 1.0.16 and Docker Desktop 4.61.0 and published by Docker, more than six months ago, -2. https://github.com/docker/model-runner/security/advisories/GHSA-m456-c56c-hh5c\n- 2026-03-30. GHSA-x2f5-332j-9xwq (CVE-2026-33990, 7.1). A malicious OCI registry could point the token exchange at an internal URL and make the runner send GET requests to host-local services. Fixed in 1.1.25 and Docker Desktop 4.67.0 and published by Docker, more than six months ago, -1. https://github.com/docker/model-runner/security/advisories/GHSA-x2f5-332j-9xwq\n\n## What we couldn't check\n\nWhat we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it.\n\n- unchecked: Docker Desktop release notes, which may carry Model Runner changes between the tagged releases, so the release count covers GitHub tags only\n- unchecked: whether the Docker Engine install publishes port 12434 on loopback only. The server source listens on every interface when `MODEL_RUNNER_PORT` is set\n- unchecked: reply times on issues. We saw comment counts on the issue list, not who replied or when\n- unchecked: Docker's SOC 2 or ISO 27001 status, which we didn't look up for this listing\n- unchecked: the first release date, and pull counts for the docker/model-runner image\n- Docker publishes no terms written for Model Runner itself. The subscription agreement and privacy policy listed are the ones that govern Docker Desktop, and the Engine plugin and `dmr` binary are under Apache-2.0 only\n- The API reference shows the Anthropic route as /anthropic/v1/messages in its table and /v1/messages in its examples. The source registers both\n\n## Weaknesses\n\n- No credential on the API. The docs say any client that can reach it, including other containers, can pull, load and run models\n- No OpenAPI file, no error reference and no rate-limit or retry guidance in the reviewed documentation\n- Two releases in the 90 days to 8 October 2026 (v1.2.7 and v1.2.8), the latest on 12 August\n- CVE-2026-28400 let an unauthenticated caller overwrite files, including the Docker Desktop VM disk, until 1.0.16 in February 2026\n- On Docker Engine the docs say model-name requests go to Docker Hub regardless of settings, and the `DO_NOT_TRACK` switch in the source is undocumented\n\n## What costs an agent a turn today\n\nThe notes we give agents before they call it. Each one is a workaround an agent shouldn't need.\n\n- Use base URL `http://localhost:12434/engines/v1` for OpenAI clients and `http://localhost:12434` for Anthropic and Ollama clients. Any API key value is accepted\n- In Docker Desktop, run `docker desktop enable model-runner --tcp 12434` first. Host-side TCP is off by default\n- From a container, call `http://model-runner.docker.internal` on Docker Desktop or `http://172.17.0.1:12434` on Docker Engine\n- Raise the context before agent work with `docker model configure --context-size \u003cn\u003e \u003cmodel\u003e`. The llama.cpp default is 4,096 tokens\n- Name models with their namespace, such as `ai/smollm2`, and expect plain-text error bodies with a 400, 404, 500 or 503 status\n\n## When it's done\n\nSend what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `\"kind\": \"dispute\"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.\n",
    "grade": "C",
    "score": 57.1,
    "assessed": "2026-10-08",
    "run": "October 2026 research run",
    "categories": [
      {
        "key": "security",
        "name": "Security \u0026 auth",
        "score": 40,
        "maxGain": 10.5,
        "reason": "Read with the tool checklist, for the local API. No credential, by design. The docs say the API is not authenticated and that any client that can reach it, including other containers on the same Docker network, can pull, load and run models. Host-side TCP is off by default in Docker Desktop and on by default on port 12434 in Docker Engine, and cross-origin requests are refused with 403 unless the origin is localhost, 127.0.0.1, 0.0.0.0 or listed in `DMR_ORIGINS` (7 of 30). No read-only mode or per-caller limit. Runtime flags pass an allowlist, engines run sandboxed on macOS and Windows and in a container on Linux, and Enhanced Container Isolation, a Docker Business control, blocks container access (6 of 20). The API returns model output, with no injection guidance in the docs (6 of 15). `docker model requests` and the Requests tab show recent requests and responses, the source keeps the last 10 per model, and `/metrics` exposes Prometheus counters. No caller identity (8 of 15). www.docker.com has a valid security.txt with a disclosure policy, SECURITY.md promises an acknowledgement within 72 hours, and two advisories were published with CVEs in 2026. No monetary bounty for this project (13 of 20).",
        "checklist": [
          "- 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option.\n- 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions.\n- 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10.\n- 0 to 15, audit logs or per-call visibility for the operator.\n- 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public.",
          "Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-security"
      },
      {
        "key": "schema",
        "name": "Schema \u0026 documentation",
        "score": 49,
        "maxGain": 8.3,
        "reason": "Read for an API. No OpenAPI or similar file was found in the repository or the docs. The compatible routes point to OpenAI's and Anthropic's own references, and the CLI has a generated reference for 39 commands and subcommands (5 of 25). docs.docker.com has llms.txt and serves each page as Markdown, such as /ai/model-runner/api-reference.md (10). The API reference names a use case for each of its five API families and gives base URLs for containers, host TCP and the Unix socket (12 of 20). Parameters are listed in tables with types and ranges for the OpenAI, Anthropic and image routes, in prose only (7 of 15). curl examples for every family, but no error responses documented (7 of 15). Dated GitHub releases with notes. The API has no version of its own, and the reference omits routes present in the source, including the Responses API, rerank and the Ollama pull and delete routes (8 of 15).",
        "checklist": [
          "APIs and MCP servers.",
          "- 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool).\n- 10, llms.txt or Markdown docs served for agents.\n- 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference.\n- 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs.\n- 0 to 15, examples and documented error responses.\n- 15, versioning and a public changelog.",
          "Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-schema"
      },
      {
        "key": "ergonomics",
        "name": "Agent ergonomics",
        "score": 58,
        "maxGain": 6.8,
        "reason": "Read for an API. Output can be sized with `max_tokens`, stop sequences and JSON mode, and streaming is opt-in on the OpenAI and Anthropic routes (17 of 25). Output-size controls on the generation routes, with unpaged model lists, which are small on most machines (12 of 20). Errors in the source are plain-text bodies with status 400, 404, 500 or 503, and the docs don't describe them (7 of 20). Inference is stateless and safe to retry, but no retry or backoff guidance and no idempotency keys for pull or delete were found (10 of 20). `model` is the only required field beyond the messages or prompt. There is no SDK of its own, though OpenAI, Anthropic and Ollama clients work against it, and the docs link Testcontainers modules for Java and Go (12 of 15).",
        "checklist": [
          "- 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries).\n- 20, pagination, filtering and output-size controls.\n- 20, actionable, documented error responses, codes and messages an agent can recover from.\n- 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations.\n- 15, sensible defaults, few required parameters, and official SDKs in at least two languages.",
          "Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-ergonomics"
      },
      {
        "key": "payments",
        "name": "Payments \u0026 pricing",
        "score": 60,
        "maxGain": 5,
        "reason": "Read with the self-hosted rule, since the API an agent calls is free software on the owner's machine. No x402, MPP or L402 in the docs or the source (0). Model Runner is Apache-2.0 with nothing to buy and no account needed for the Docker Engine plugin or the `dmr` binary, so 20, 20 and 20 on the last three lines. One qualification applies to the Docker Desktop route. Docker's subscription agreement limits free Desktop use to non-commercial open-source projects and businesses with fewer than 250 employees and under US $10,000,000 in annual revenue, and paid plans run from $11 to $24 per user a month.",
        "checklist": [
          "The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/).",
          "- 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which.\n- 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for \"contact sales\" or prices behind a login.\n- 20, a free tier or trial that doesn't need a card.\n- 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API).",
          "Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied.",
          "Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-payments"
      },
      {
        "key": "maintenance",
        "name": "Maintenance \u0026 community",
        "score": 55,
        "maxGain": 3.9,
        "reason": "v1.2.8 was released on 12 August 2026, 57 days before the check (20 of 30). Two releases in the 90 days to 8 October 2026, v1.2.7 on 11 August and v1.2.8, against 30 or so between March and June (0 of 20). 45 commits on main since 10 July, the newest on 8 October, and each of the 20 newest open issues had at least one comment, with 41 issues and 25 pull requests open (17 of 25). No SDK of its own. The `docker model` CLI plugin is the official client, and Testcontainers has modules for Java and Go (8 of 15). Dependabot runs weekly on Go modules and GitHub Actions, actions are pinned by commit, a script bumps llama.cpp, and CI passes on main (10).",
        "checklist": [
          "- 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older.\n- 20, at least three releases or dated changelog entries in the last 90 days.\n- 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15.\n- 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models).\n- 10, package health, current dependencies and CI.",
          "Models are read for deprecation notice periods and model churn rather than release counts."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-maintenance"
      },
      {
        "key": "reliability",
        "name": "Reliability",
        "score": 85,
        "maxGain": 3,
        "reason": "Read with the local-software lines, since the API runs on the owner's machine. Bundled with Docker Desktop, installed as `docker-model-plugin` from Docker's apt and dnf repositories, and shipped as a standalone `dmr` binary through Homebrew and winget, with platform, GPU and driver requirements stated in the docs (20). The CI workflow runs lint, race-detector tests and builds on pushes and pull requests to main, with separate end-to-end, integration and daily check workflows, and the ten most recent CI runs on main passed on 8 October 2026 (25). 41 open issues and 25 open pull requests. Each of the 20 newest open issues had at least one comment, but several are regressions or failures still open, including HTTP 500 on sequential tool calls (#1063), a Windows GPU regression after Docker Desktop 4.82.0 (#1054) and a context-size setting applied nondeterministically (#1025) (17 of 25). Semver tags with notes on each GitHub release, but no changelog file and no breaking-change section in the notes we read (8 of 15). Version 1.2.8 (15). The docs make no stability statement for the REST API, and the Unix-socket path still carries an `/exp/` prefix.",
        "checklist": [
          "Hosted APIs, MCP servers, models and platforms.",
          "- 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own).\n- 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so.\n- 15, rate limits documented with numbers.\n- 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved.\n- 10, an SLA published for any paid tier.\n- 10, the surface agents use is generally available, not beta or preview.",
          "Local packages, SDKs, frameworks and stdio MCP servers.",
          "- 20, installs from an official package with supported runtimes stated.\n- 25, a public CI and test suite, passing on the default branch.\n- 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered).\n- 15, semver discipline and breaking changes called out in a changelog.\n- 15, version 1.0 or later, or declared stable.",
          "Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-reliability"
      },
      {
        "key": "transparency",
        "name": "Transparency \u0026 trust",
        "score": 73,
        "maxGain": 2.4,
        "reason": "The editorial half. Apache-2.0 for the server, the CLI and the `dmr` binary in a public repository. Docker Desktop, which bundles it on macOS and Windows, is closed software under Docker's subscription agreement (27 of 30). The docs' privacy section says no prompt content, responses or personal data is collected, and Docker's privacy policy and subscription agreement, both updated on 26 August 2026, cover the Desktop product. No retention period specific to Model Runner was found (18 of 30). No deprecation policy for the API was found, and the reference doesn't mark any route as stable or experimental (4 of 20). Telemetry is disclosed with a link to the source. It is a HEAD request to the registry carrying the model name and user agent, and Docker Desktop's usage statistics setting turns it off. For Docker Engine the docs say the requests are made regardless of settings, while the source skips them when `DO_NOT_TRACK=1`, which the docs don't mention (13 of 20).",
        "blend": "editorial 62, provenance 84",
        "checklist": [
          "- 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms.\n- 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors).\n- 0 to 20, a deprecation policy or notices with dates.\n- 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted).",
          "The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two."
        ],
        "checklistUrl": "https://www.anchorterminal.com/benchmark/#checklist-transparency"
      }
    ],
    "provenance": [
      {
        "label": "Terms of service",
        "value": "read, states 7 of the 7 things a reader expects, and has 2 clauses that cost points",
        "points": 6,
        "max": 10
      },
      {
        "label": "Status page",
        "value": "not found",
        "points": 0,
        "max": 10
      }
    ],
    "deductions": [
      "2026-02-27. GHSA-m456-c56c-hh5c (CVE-2026-28400, 7.5). The unauthenticated `/engines/_configure` route accepted arbitrary runtime flags, so a caller, including a container on Docker Desktop, could overwrite files the runner could reach, the Desktop VM disk among them. Fixed in Model Runner 1.0.16 and Docker Desktop 4.61.0 and published by Docker, more than six months ago, -2. https://github.com/docker/model-runner/security/advisories/GHSA-m456-c56c-hh5c",
      "2026-03-30. GHSA-x2f5-332j-9xwq (CVE-2026-33990, 7.1). A malicious OCI registry could point the token exchange at an internal URL and make the runner send GET requests to host-local services. Fixed in 1.1.25 and Docker Desktop 4.67.0 and published by Docker, more than six months ago, -1. https://github.com/docker/model-runner/security/advisories/GHSA-x2f5-332j-9xwq"
    ],
    "unchecked": [
      "unchecked: Docker Desktop release notes, which may carry Model Runner changes between the tagged releases, so the release count covers GitHub tags only",
      "unchecked: whether the Docker Engine install publishes port 12434 on loopback only. The server source listens on every interface when `MODEL_RUNNER_PORT` is set",
      "unchecked: reply times on issues. We saw comment counts on the issue list, not who replied or when",
      "unchecked: Docker's SOC 2 or ISO 27001 status, which we didn't look up for this listing",
      "unchecked: the first release date, and pull counts for the docker/model-runner image",
      "Docker publishes no terms written for Model Runner itself. The subscription agreement and privacy policy listed are the ones that govern Docker Desktop, and the Engine plugin and `dmr` binary are under Apache-2.0 only",
      "The API reference shows the Anthropic route as /anthropic/v1/messages in its table and /v1/messages in its examples. The source registers both"
    ],
    "weaknesses": [
      "No credential on the API. The docs say any client that can reach it, including other containers, can pull, load and run models",
      "No OpenAPI file, no error reference and no rate-limit or retry guidance in the reviewed documentation",
      "Two releases in the 90 days to 8 October 2026 (v1.2.7 and v1.2.8), the latest on 12 August",
      "CVE-2026-28400 let an unauthenticated caller overwrite files, including the Docker Desktop VM disk, until 1.0.16 in February 2026",
      "On Docker Engine the docs say model-name requests go to Docker Hub regardless of settings, and the `DO_NOT_TRACK` switch in the source is undocumented"
    ],
    "agentNotes": [
      "Use base URL `http://localhost:12434/engines/v1` for OpenAI clients and `http://localhost:12434` for Anthropic and Ollama clients. Any API key value is accepted",
      "In Docker Desktop, run `docker desktop enable model-runner --tcp 12434` first. Host-side TCP is off by default",
      "From a container, call `http://model-runner.docker.internal` on Docker Desktop or `http://172.17.0.1:12434` on Docker Engine",
      "Raise the context before agent work with `docker model configure --context-size \u003cn\u003e \u003cmodel\u003e`. The llama.cpp default is 4,096 tokens",
      "Name models with their namespace, such as `ai/smollm2`, and expect plain-text error bodies with a 400, 404, 500 or 503 status"
    ],
    "recheck": "https://www.anchorterminal.com/builders/#disputes"
  },
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-08",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  }
}
