Head to head · Local inference · October 2026 research run

Foundry Local vs vLLM

Foundry Local scores 60.5 (C) on agent readiness against vLLM's 57.7 (C), and leads in 2 of 7 scored categories. vLLM leads on schema & documentation and security & auth. Both do local inference.

Best local AI models and assistants · All 184 local ai comparisons

Which one, for what

Foundry Local C

Good for An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.

Ahead on

  • Reliability, 68 against 62
  • Transparency & trust, 72 against 67

Also in its favour

  • No incidents deducted, where vLLM loses 6 points for them

Watch for

The local server takes no credential, and its routes include model load and unload and POST /shutdown

vLLM C

Good for An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.

Ahead on

  • Schema & documentation, 68 against 53
  • Security & auth, 50 against 39

Watch for

--api-key guards only the /v1, /v2, /inference and /cohere prefixes. /invocations, /pooling, /classify, /score, /rerank, /pause and /update_weights answer without it

Score by category

CategoryWeight this runFoundry LocalvLLMEdge
Reliability16%206862Foundry Local +6
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25368vLLM +15
Agent ergonomics13%16.26164vLLM +3
Security & auth14%17.53950vLLM +11
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88888even
Transparency & trust7%8.87267Foundry Local +5
Negative events≤150-6
Total60.5 · C57.7 · C

Facts side by side

FactFoundry LocalvLLM
KindSDK + MCPHTTP API
VendorMicrosoftvLLM project (PyTorch Foundation)
Hosted endpointno (local only)no (local only)
TransportsHTTPHTTP
AuthNoneNone
PricingFreeFree
x402nono
LicenceMIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its ownApache-2.0
Read-only variant documentednono
llms.txtnono
Last release2026-09-292026-10-02
Terms last updatedno document linkedno document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train models
Terms restrict automated access
Terms restrict benchmarking
Terms or service can change without notice
Arbitration or class-action waiver
Popularity2.6k stars, 105k npm/wk, 47k PyPI/wk93k stars

Verdicts

Foundry Local

The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.

vLLM

Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so /invocations and control routes such as /pause answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.

Before you call either

Foundry Local

  1. Read the server URL from manager.urls[0] or foundry server status. The port is dynamic unless the owner sets web.urls or foundry server start --port
  2. Send the model ID that GET /v1/models returns, not the alias. The alias resolves to a hardware-specific variant
  3. Check supportsToolCalling before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN
  4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step
  5. Ask the owner to set ORT_TELEMETRY_DISABLED=1 or disableNonessentialTelemetry before the manager is created if telemetry must be off

vLLM

  1. Put a reverse proxy that allowlists routes in front of the server. --api-key leaves /invocations and the control routes open
  2. Pass --host 127.0.0.1 for single-machine use. With no --host the server listens on every interface
  3. Set VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1 before starting if nothing should be sent to stats.vllm.ai
  4. Start with --enable-auto-tool-choice and the --tool-call-parser for the model before sending tools. Tool calling is off without them
  5. Send max_tokens on every request, and read the breaking changes section of the release notes before upgrading a minor version

Questions

Which is better for AI agents, Foundry Local or vLLM?

Foundry Local scores 60.5 (C) on agent readiness against vLLM's 57.7 (C), and leads in 2 of 7 scored categories. vLLM leads on schema & documentation and security & auth.

Can an agent call Foundry Local and vLLM without installing anything?

No hosted endpoint is listed for Foundry Local. No hosted endpoint is listed for vLLM.

Are Foundry Local and vLLM open source?

Yes. Foundry Local is open source (MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own). vLLM is open source (Apache-2.0).

Other comparisons with Foundry Local or vLLM

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.