# Fix list: Foundry Local From Anchor Terminal's listing at https://www.anchorterminal.com/tools/foundry-local, the October 2026 research run, assessed 8 October 2026. Grade C, 60.5 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on Foundry Local: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Security & auth, 39 out of 100, up to 10.7 more on the total Why it scored 39: Read with the tool checklist, for the local server. No credential exists. The server is off until the application or the CLI starts it, binds to 127.0.0.1 by default, and the docs tell clients to set auth to None (8 of 30). No read-only mode. Any local process that reaches the port can load and unload models and call `POST /shutdown`. Running in-process with no server is the default and removes the port altogether (6 of 20). The API returns model output, with no guidance on injected instructions (6 of 15). `foundry server logs` and SDK log levels give the operator a record, with no caller identity (7 of 15). SECURITY.md sends reports to MSRC and promises a reply within 24 hours, GitHub Actions are pinned to commit hashes, and no advisory or CVE was found. The security.txt on www.microsoft.com expired on 23 September 2026, and open issue #1123 says binaries downloaded at runtime rely on transport integrity alone (12 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 2. Schema & documentation, 53 out of 100, up to 7.6 more on the total Why it scored 53: Read with the API lines, for the local server and the SDKs. No OpenAPI file in the repository or the docs. The server follows OpenAI's shapes and the SDKs are typed (8 of 25). Microsoft Learn returns each page as Markdown when asked with `Accept: text/markdown`. No llms.txt on learn.microsoft.com or foundrylocal.ai (6 of 10). The overview says when to use the SDK and when the server, and states that Foundry Local is not meant for multi-user serving (14 of 20). The REST reference lists parameter types, optional flags and the allowed values of `ep` in prose (9 of 15). 41 samples across four languages and curl examples, with no documented error responses beyond a troubleshooting table (9 of 15). Dated release notes on GitHub. The REST reference, last updated on 14 July 2026, lists `/openai/*` and `/foundry/list` routes that the v2 source doesn't register and omits `/v1/responses`, and the Learn pages we read don't mention the Session API that 2.0.1 made the main interface and still give install commands for the `-winml` packages that release removed (7 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 3. Reliability, 68 out of 100, up to 6.4 more on the total Why it scored 68: Read with the local-software lines, since Foundry Local runs in the owner's application or as a local daemon. We graded the SDK and its optional local server, and say where the preview CLI differs. Official packages on npm, PyPI, NuGet and crates.io, the CLI through winget and Homebrew, with Node.js 20, Python 3.11 to 3.14, .NET 8 and Rust 1.70 stated (20). The repository holds test suites for the C++ runtime and each SDK, but the only public GitHub Actions workflow is a build check for samples. Build and test run in Azure Pipelines whose results we couldn't see, and the pipeline file says its push trigger is switched off (12 of 25). 76 open issues against 362 closed. Open reports include a Linux ARM64 segmentation fault (#1182, 8 October), HTTP 500 on Qwen3.5 CUDA models (#1179), memory kept after unload (#1079) and several GPU and NPU detection failures from September with no reply (13 of 25). Versioned releases with notes on GitHub, and v2.0.1 has a breaking-changes section. There is no changelog file, and the CLI's REST reference warns of breaking changes without notice (10 of 15). SDK 2.1.0. The PyPI package still carries an Alpha classifier and the CLI is a public preview (13 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 4. Agent ergonomics, 61 out of 100, up to 6.3 more on the total Why it scored 61: Read with the API lines. Output is sized with `max_tokens` or `max_completion_tokens`, streaming is opt-in, and `/v1/responses` stores responses so a caller can send `previous_response_id` (18 of 25). `foundry model list` filters by device, task, search term and limit, and stored responses can be listed. Model lists over HTTP aren't paged (12 of 20). Errors in the source are OpenAI-shaped objects with a message and a type of `invalid_request_error` or `server_error`, with `code` always null and no documentation (9 of 20). Inference calls are stateless and safe to repeat. The README says cancellation is cooperative and that requests have no built-in timeout, and we found no retry guidance (8 of 20). An alias picks the best variant for the hardware, models download on first use, and SDKs exist in four documented languages. The server's port is dynamic by default, so a caller has to discover it (14 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 5. Payments & pricing, 60 out of 100, up to 5 more on the total Why it scored 60: Read with the self-hosted rule, since everything an agent calls runs on the owner's machine. No x402, MPP or L402 (0). The SDK is MIT, the CLI is a free download, and no account, card or Azure subscription is needed, so 20, 20 and 20 on the last three lines. The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 6. Transparency & trust, 72 out of 100, up to 2.5 more on the total Made of editorial 63, provenance 80. Why it scored 72: The editorial half. The SDKs and the v2 native runtime are MIT in a public repository. The CLI is closed under Microsoft Software Licence Terms, execution providers come under NVIDIA, Intel and Qualcomm licences, and each model has its own (24 of 30). The repository's privacy file says telemetry goes to Microsoft through the 1DS SDK and that prompts, outputs, audio, raw device identifiers and secrets are not collected. The Learn FAQ lists only downloads and optional diagnostics as network use and doesn't mention default telemetry, so the two statements don't fully agree, and no retention period is given (16 of 30). The v2.0.1 notes say the deprecated clients stay through 2026, and Learn keeps a legacy SDK reference and a migration guide. There is no written deprecation policy (9 of 20). Telemetry is disclosed in the README and is on by default. A config option or `ORT_TELEMETRY_DISABLED=1` turns off non-essential telemetry, and what counts as essential isn't stated (14 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Status page: not found (0 of 10) - security.txt: published but past its Expires date (5 of 10) ## 7. Maintenance & community, 88 out of 100, up to 1.1 more on the total Why it scored 88: SDK 2.1.0 reached the registries on 29 September 2026 and CLI preview 0.11.0 followed on 7 October (30). Six releases in the 90 days to 8 October (20). 132 commits on main since 10 July, the newest on 7 October, and 362 issues closed against 76 open. Many of the newest open issues had no comment when we read them, and SUPPORT.md is still Microsoft's unedited template (15 of 25). Current SDKs at the same version on npm, PyPI, NuGet and crates.io (15). Dependabot is configured, native dependency versions are pinned in one file that the build validates, and actions are pinned. We couldn't see CI results (8 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - Unchecked: whether the preview CLI's daemon adds the `/openai/*` and `/foundry/list` routes of the REST reference on top of the v2 runtime's routes. The CLI is closed source and we didn't run it - Unchecked: CI results. Build and test run in Azure Pipelines, which we couldn't read, and api.github.com refused our requests for rate limiting - Unchecked: the Microsoft Privacy Statement, which answered 403, so retention and sharing terms for the telemetry were not read - Unchecked: how many models the catalogue holds. foundrylocal.ai/models is drawn by script - What the telemetry events contain and which of them count as essential. The privacy file doesn't list them - Whether the local server checks the Origin or Host header. We found no such check in the service source but did not test it - The star count (2.6k) and issue counts come from GitHub's web pages as read through a summarising fetcher ## Weaknesses - The local server takes no credential, and its routes include model load and unload and `POST /shutdown` - The REST reference on Microsoft Learn lists `/openai/*` and `/foundry/list` routes that the v2 runtime source doesn't register, and doesn't cover `/v1/responses` - Telemetry is on by default through Microsoft's 1DS SDK. The opt-out covers non-essential telemetry only, and the Learn FAQ doesn't mention it - The CLI is a closed-source public preview, and its REST reference warns of breaking changes without notice - 76 open issues on 8 October 2026, among them a Linux ARM64 segmentation fault (#1182) and several unanswered GPU and NPU detection reports ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Read the server URL from `manager.urls[0]` or `foundry server status`. The port is dynamic unless the owner sets `web.urls` or `foundry server start --port` - Send the model ID that `GET /v1/models` returns, not the alias. The alias resolves to a hardware-specific variant - Check `supportsToolCalling` before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN - Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step - Ask the owner to set `ORT_TELEMETRY_DISABLED=1` or `disableNonessentialTelemetry` before the manager is created if telemetry must be off ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.