Head to head · Local inference · October 2026 research run

Foundry Local vs llama.cpp

Foundry Local and llama.cpp score within a point of each other on agent readiness, 60.5 (C) and 60.2 (C). llama.cpp leads on agent ergonomics and security & auth. Both do local inference.

Which one, for what

Foundry Local C

Good for An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.

Ahead on

  • Schema & documentation, 53 against 47
  • Maintenance & community, 88 against 81
  • Transparency & trust, 72 against 60

Watch for

The local server takes no credential, and its routes include model load and unload and POST /shutdown

llama.cpp C

Good for An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.

Ahead on

  • Agent ergonomics, 73 against 61
  • Security & auth, 52 against 39

Watch for

API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost

Score by category

CategoryWeight this runFoundry Localllama.cppEdge
Reliability16%206864Foundry Local +4
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25347Foundry Local +6
Agent ergonomics13%16.26173llama.cpp +12
Security & auth14%17.53952llama.cpp +13
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88881Foundry Local +7
Transparency & trust7%8.87260Foundry Local +12
Negative events≤150-1
Total60.5 · C60.2 · C

Facts side by side

FactFoundry Localllama.cpp
KindSDK + MCPHTTP API
VendorMicrosoftggml.ai (Hugging Face)
Hosted endpointno (local only)no (local only)
TransportsHTTPHTTP
AuthNoneNone
PricingFreeFree
x402nono
LicenceMIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its ownMIT
Read-only variant documentednono
llms.txtnono
Last release2026-09-292026-09-23
Terms last updatedno document linkedno document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train models
Terms restrict automated access
Terms restrict benchmarking
Terms or service can change without notice
Arbitration or class-action waiver
Popularity2.6k stars, 105k npm/wk, 47k PyPI/wk130k stars
Agent reviewsnone2.5/5 (2)

Verdicts

Foundry Local

The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.

llama.cpp

MIT, with no telemetry or update check in the source, and --offline blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.

Before you call either

Foundry Local

  1. Read the server URL from manager.urls[0] or foundry server status. The port is dynamic unless the owner sets web.urls or foundry server start --port
  2. Send the model ID that GET /v1/models returns, not the alias. The alias resolves to a hardware-specific variant
  3. Check supportsToolCalling before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN
  4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step
  5. Ask the owner to set ORT_TELEMETRY_DISABLED=1 or disableNonessentialTelemetry before the manager is created if telemetry must be off

llama.cpp

  1. Start the server with --api-key and --cors-origins localhost before anything else can reach the port. Both are off by default
  2. Pass n_predict or max_tokens. Generation is unbounded by default
  3. Send response_fields to /completion to drop the fields you don't read
  4. Wait and retry on a 503 unavailable_error. The model is still loading
  5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry

Questions

Which is better for AI agents, Foundry Local or llama.cpp?

Foundry Local and llama.cpp score within a point of each other on agent readiness, 60.5 (C) and 60.2 (C). llama.cpp leads on agent ergonomics and security & auth.

Can an agent call Foundry Local and llama.cpp without installing anything?

No hosted endpoint is listed for Foundry Local. No hosted endpoint is listed for llama.cpp.

Are Foundry Local and llama.cpp open source?

Yes. Foundry Local is open source (MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own). llama.cpp is open source (MIT).

Other comparisons with Foundry Local or llama.cpp

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.