Head to head · Local inference · October 2026 research run

Foundry Local vs KoboldCpp

Foundry Local and KoboldCpp score within a point of each other on agent readiness, 60.5 (C) and 60.5 (C). KoboldCpp leads on schema & documentation. Both do local inference.

Which one, for what

Foundry Local C

Good for An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.

Ahead on

  • Maintenance & community, 88 against 82
  • Transparency & trust, 72 against 49

Watch for

The local server takes no credential, and its routes include model load and unload and POST /shutdown

KoboldCpp C

Good for An owner who wants text, image, speech and music models behind one executable with a writing and roleplay interface, and clients that speak the KoboldAI, OpenAI, Ollama or Anthropic formats.

Ahead on

  • Schema & documentation, 68 against 53

Watch for

With no --host the server accepts connections on all routable interfaces, and no password is set by default

Score by category

CategoryWeight this runFoundry LocalKoboldCppEdge
Reliability16%206868even
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25368KoboldCpp +15
Agent ergonomics13%16.26163KoboldCpp +2
Security & auth14%17.53938Foundry Local +1
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88882Foundry Local +6
Transparency & trust7%8.87249Foundry Local +23
Negative events≤1500
Total60.5 · C60.5 · C

Facts side by side

FactFoundry LocalKoboldCpp
KindSDK + MCPHTTP API
VendorMicrosoftLostRuins (Concedo)
Hosted endpointno (local only)no (local only)
TransportsHTTPHTTP
AuthNoneNone
PricingFreeFree
x402nono
LicenceMIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its ownAGPL-3.0 for KoboldCpp and KoboldAI Lite. The bundled GGML, llama.cpp and stable-diffusion.cpp code stays under MIT
Read-only variant documentednono
llms.txtnoyes
Last release2026-09-292026-09-27
Terms last updatedno document linkedno document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train models
Terms restrict automated access
Terms restrict benchmarking
Terms or service can change without notice
Arbitration or class-action waiver
Popularity2.6k stars, 105k npm/wk, 47k PyPI/wk12k stars

Verdicts

Foundry Local

The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.

KoboldCpp

One file runs text, image, speech and music models behind a published OpenAPI 3.0.3 document, with eight releases in 90 days. The server listens on every interface with no password by default, and --password leaves the image routes open.

Before you call either

Foundry Local

  1. Read the server URL from manager.urls[0] or foundry server status. The port is dynamic unless the owner sets web.urls or foundry server start --port
  2. Send the model ID that GET /v1/models returns, not the alias. The alias resolves to a hardware-specific variant
  3. Check supportsToolCalling before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN
  4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step
  5. Ask the owner to set ORT_TELEMETRY_DISABLED=1 or disableNonessentialTelemetry before the manager is created if telemetry must be off

KoboldCpp

  1. Start with --host 127.0.0.1 and --password. The default listens on every interface with no key
  2. Send the password as Authorization: Bearer <password>. It is not read from the query string
  3. Treat 503 as both busy and rate limited. The server never sends 429 or Retry-After, and the wait in seconds is in detail.msg
  4. Pass max_length or max_tokens. The default is 2,048 tokens unless --defaultgenamt changes it
  5. Send a genkey with each generation so /api/extra/generate/check and /api/extra/abort act on your request and not another caller's

Questions

Which is better for AI agents, Foundry Local or KoboldCpp?

Foundry Local and KoboldCpp score within a point of each other on agent readiness, 60.5 (C) and 60.5 (C). KoboldCpp leads on schema & documentation.

Can an agent call Foundry Local and KoboldCpp without installing anything?

No hosted endpoint is listed for Foundry Local. No hosted endpoint is listed for KoboldCpp.

Are Foundry Local and KoboldCpp open source?

Yes. Foundry Local is open source (MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own). KoboldCpp is open source (AGPL-3.0 for KoboldCpp and KoboldAI Lite. The bundled GGML, llama.cpp and stable-diffusion.cpp code stays under MIT).

Other comparisons with Foundry Local or KoboldCpp

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.