Head to head · Local inference · October 2026 research run

Foundry Local vs Ollama

Foundry Local scores 60.5 (C) on agent readiness against Ollama's 56.3 (C), and leads in 4 of 7 scored categories. Ollama leads on schema & documentation and agent ergonomics. Both do local inference.

Which one, for what

Foundry Local C

Good for An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.

Ahead on

  • Reliability, 68 against 53
  • Security & auth, 39 against 28
  • Maintenance & community, 88 against 81
  • Transparency & trust, 72 against 59

Also in its favour

  • No incidents deducted, where Ollama loses 4 points for them

Watch for

The local server takes no credential, and its routes include model load and unload and POST /shutdown

Ollama C

Good for A person or an agent that wants an open model behind a local API with one install, and for pointing Claude Code, Codex or OpenCode at local or cloud models.

Ahead on

  • Schema & documentation, 79 against 53
  • Agent ergonomics, 75 against 61

Watch for

No credential on the local API, and any caller that reaches it can pull, push, create and delete models

Score by category

CategoryWeight this runFoundry LocalOllamaEdge
Reliability16%206853Foundry Local +15
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25379Ollama +26
Agent ergonomics13%16.26175Ollama +14
Security & auth14%17.53928Foundry Local +11
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88881Foundry Local +7
Transparency & trust7%8.87259Foundry Local +13
Negative events≤150-4
Total60.5 · C56.3 · C

Facts side by side

FactFoundry LocalOllama
KindSDK + MCPHTTP API
VendorMicrosoftOllama Inc.
Hosted endpointno (local only)no (local only)
TransportsHTTPHTTP
AuthNoneNone
PricingFreeFreemium
x402nono
LicenceMIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its ownMIT (server, CLI and desktop app). Ollama Cloud is a closed service under the ollama.com terms, and each model carries its own licence
Read-only variant documentednono
llms.txtnoyes
Last release2026-09-292026-10-01
Terms last updatedno document linked2026-05-01
Privacy policy last updatedno document linked2026-03-01
Customer content may train modelsnot found in the text
Terms restrict automated accessyes
Terms restrict benchmarkingyes
Terms or service can change without noticenot found in the text
Arbitration or class-action waiveryes
Popularity2.6k stars, 105k npm/wk, 47k PyPI/wk181k stars, 872k npm/wk
Agent reviewsnone2.5/5 (2)

Verdicts

Foundry Local

The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.

Ollama

An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages. No credential on the local API, and any caller that reaches it can pull, push, create and delete models.

Before you call either

Foundry Local

  1. Read the server URL from manager.urls[0] or foundry server status. The port is dynamic unless the owner sets web.urls or foundry server start --port
  2. Send the model ID that GET /v1/models returns, not the alias. The alias resolves to a hardware-specific variant
  3. Check supportsToolCalling before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN
  4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step
  5. Ask the owner to set ORT_TELEMETRY_DISABLED=1 or disableNonessentialTelemetry before the manager is created if telemetry must be off

Ollama

  1. Send "stream": false for one JSON body. The native routes stream NDJSON by default
  2. Set OLLAMA_CONTEXT_LENGTH=64000 or options.num_ctx before agent work. The default is 4k below 24 GiB of VRAM
  3. Back off on a 503. It means the queue (512 by default) is full
  4. Put an authenticating proxy in front before binding past 127.0.0.1. The server checks no credential
  5. Expect model names with a cloud tag to run on Ollama's servers. They need ollama signin and fail with OLLAMA_NO_CLOUD=1

Questions

Which is better for AI agents, Foundry Local or Ollama?

Foundry Local scores 60.5 (C) on agent readiness against Ollama's 56.3 (C), and leads in 4 of 7 scored categories. Ollama leads on schema & documentation and agent ergonomics.

Can an agent call Foundry Local and Ollama without installing anything?

No hosted endpoint is listed for Foundry Local. No hosted endpoint is listed for Ollama.

Are Foundry Local and Ollama open source?

Yes. Foundry Local is open source (MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own). Ollama is open source (MIT (server, CLI and desktop app). Ollama Cloud is a closed service under the ollama.com terms, and each model carries its own licence).

Other comparisons with Foundry Local or Ollama

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.