# Foundry Local vs vLLM > Foundry Local scores 60.5 (C) to vLLM's 57.7 (C) for local inference. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/foundry-local-vs-vllm - Markdown: https://www.anchorterminal.com/compare/foundry-local-vs-vllm.md (~2,650 tokens) - Slim: https://www.anchorterminal.com/compare/foundry-local-vs-vllm.min.md (~580 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/foundry-local-vs-vllm.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 Foundry Local scores 60.5 (C) on agent readiness against vLLM's 57.7 (C), and leads in 2 of 7 scored categories. vLLM leads on schema & documentation and security & auth. Both do local inference. - Foundry Local: grade C, 60.5/100, rank #509 of 950. Markdown https://www.anchorterminal.com/tools/foundry-local.md · JSON https://www.anchorterminal.com/api/v1/tools/foundry-local.json - vLLM: grade C, 57.7/100, rank #600 of 950. Markdown https://www.anchorterminal.com/tools/vllm.md · JSON https://www.anchorterminal.com/api/v1/tools/vllm.json - Best local AI models and assistants: https://www.anchorterminal.com/best/local-ai/index.md - All 184 local ai comparisons: https://www.anchorterminal.com/compare/local-ai/index.md ## Which one, for what ### Foundry Local (C) Good for: An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows. Ahead on: - Reliability, 68 against 62 - Transparency & trust, 72 against 67 Also in its favour: - No incidents deducted, where vLLM loses 6 points for them Watch for: The local server takes no credential, and its routes include model load and unload and `POST /shutdown` ### vLLM (C) Good for: An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes. Ahead on: - Schema & documentation, 68 against 53 - Security & auth, 50 against 39 Watch for: `--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it ## Score by category | Category | Weight | Foundry Local | vLLM | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 68 | 62 | Foundry Local +6 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 53 | 68 | vLLM +15 | | Agent ergonomics | 13% (16.2 this run) | 61 | 64 | vLLM +3 | | Security & auth | 14% (17.5 this run) | 39 | 50 | vLLM +11 | | Payments & pricing | 10% (12.5 this run) | 60 | 60 | even | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 88 | 88 | even | | Transparency & trust | 7% (8.8 this run) | 72 | 67 | Foundry Local +5 | | Negative events | ≤15 | 0 | -6 | | | **Total** | | **60.5 · C** | **57.7 · C** | | ## Facts side by side | Fact | Foundry Local | vLLM | | --- | --- | --- | | Kind | SDK + MCP | HTTP API | | Vendor | Microsoft | vLLM project (PyTorch Foundation) | | Hosted endpoint | no (local only) | no (local only) | | Transports | HTTP | HTTP | | Auth | None | None | | Pricing | Free | Free | | x402 | no | no | | Licence | MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own | Apache-2.0 | | Read-only variant documented | no | no | | llms.txt | no | no | | Last release | 2026-09-29 | 2026-10-02 | | Terms last updated | no document linked | no document linked | | Privacy policy last updated | no document linked | no document linked | | Customer content may train models | | | | Terms restrict automated access | | | | Terms restrict benchmarking | | | | Terms or service can change without notice | | | | Arbitration or class-action waiver | | | | Popularity | 2.6k stars, 105k npm/wk, 47k PyPI/wk | 93k stars | ## Verdicts **Foundry Local.** The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out. **vLLM.** Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026. ## Before you call either ### Foundry Local 1. Read the server URL from `manager.urls[0]` or `foundry server status`. The port is dynamic unless the owner sets `web.urls` or `foundry server start --port` 2. Send the model ID that `GET /v1/models` returns, not the alias. The alias resolves to a hardware-specific variant 3. Check `supportsToolCalling` before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN 4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step 5. Ask the owner to set `ORT_TELEMETRY_DISABLED=1` or `disableNonessentialTelemetry` before the manager is created if telemetry must be off ### vLLM 1. Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open 2. Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface 3. Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai 4. Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them 5. Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version ## Questions ### Which is better for AI agents, Foundry Local or vLLM? Foundry Local scores 60.5 (C) on agent readiness against vLLM's 57.7 (C), and leads in 2 of 7 scored categories. vLLM leads on schema & documentation and security & auth. ### Can an agent call Foundry Local and vLLM without installing anything? No hosted endpoint is listed for Foundry Local. No hosted endpoint is listed for vLLM. ### Are Foundry Local and vLLM open source? Yes. Foundry Local is open source (MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own). vLLM is open source (Apache-2.0). ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/foundry-local-vs-vllm.json, and with the fewest tokens: https://www.anchorterminal.com/compare/foundry-local-vs-vllm.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "foundry-local", "b": "vllm"}`. From a terminal: `anchor compare foundry-local vllm` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/foundry-local.json and https://www.anchorterminal.com/api/v1/tools/vllm.json ## Other comparisons with Foundry Local or vLLM - [AnythingLLM vs Foundry Local](https://www.anchorterminal.com/compare/anythingllm-vs-foundry-local.md) - [AnythingLLM vs vLLM](https://www.anchorterminal.com/compare/anythingllm-vs-vllm.md) - [Docker Model Runner vs Foundry Local](https://www.anchorterminal.com/compare/docker-model-runner-vs-foundry-local.md) - [Docker Model Runner vs vLLM](https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.md) - [Foundry Local vs Core](https://www.anchorterminal.com/compare/foundry-local-vs-ghost-core.md) - [Foundry Local vs GPT4All](https://www.anchorterminal.com/compare/foundry-local-vs-gpt4all.md) - [Foundry Local vs Jan](https://www.anchorterminal.com/compare/foundry-local-vs-jan.md) - [Foundry Local vs Khoj](https://www.anchorterminal.com/compare/foundry-local-vs-khoj.md) - [Foundry Local vs KoboldCpp](https://www.anchorterminal.com/compare/foundry-local-vs-koboldcpp.md) - [Foundry Local vs Lemonade](https://www.anchorterminal.com/compare/foundry-local-vs-lemonade.md) - [Foundry Local vs llama.cpp](https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp.md) - [Foundry Local vs LM Studio](https://www.anchorterminal.com/compare/foundry-local-vs-lm-studio.md) - [Foundry Local vs LocalAI](https://www.anchorterminal.com/compare/foundry-local-vs-localai.md) - [Foundry Local vs MLX LM](https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm.md) - [Foundry Local vs Ollama](https://www.anchorterminal.com/compare/foundry-local-vs-ollama.md) - [Foundry Local vs Open WebUI](https://www.anchorterminal.com/compare/foundry-local-vs-open-webui.md) - [Foundry Local vs screenpipe](https://www.anchorterminal.com/compare/foundry-local-vs-screenpipe.md) - [Foundry Local vs TextGen](https://www.anchorterminal.com/compare/foundry-local-vs-text-generation-webui.md) - [Core vs vLLM](https://www.anchorterminal.com/compare/ghost-core-vs-vllm.md) - [GPT4All vs vLLM](https://www.anchorterminal.com/compare/gpt4all-vs-vllm.md) - [Jan vs vLLM](https://www.anchorterminal.com/compare/jan-vs-vllm.md) - [Khoj vs vLLM](https://www.anchorterminal.com/compare/khoj-vs-vllm.md) - [KoboldCpp vs vLLM](https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.md) - [Lemonade vs vLLM](https://www.anchorterminal.com/compare/lemonade-vs-vllm.md) - [llama.cpp vs vLLM](https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.md) - [LM Studio vs vLLM](https://www.anchorterminal.com/compare/lm-studio-vs-vllm.md) - [LocalAI vs vLLM](https://www.anchorterminal.com/compare/localai-vs-vllm.md) - [MLX LM vs vLLM](https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.md) - [Ollama vs vLLM](https://www.anchorterminal.com/compare/ollama-vs-vllm.md) - [Open WebUI vs vLLM](https://www.anchorterminal.com/compare/open-webui-vs-vllm.md) - [screenpipe vs vLLM](https://www.anchorterminal.com/compare/screenpipe-vs-vllm.md) - [TextGen vs vLLM](https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.md) - [Foundry Local vs Underdog](https://www.anchorterminal.com/compare/foundry-local-vs-underdog.md) - [Underdog vs vLLM](https://www.anchorterminal.com/compare/underdog-vs-vllm.md)