# llama.cpp vs vLLM > llama.cpp scores 60.2 (C) to vLLM's 57.7 (C) for local inference. Prices, MCP, x402, uptime and agent notes side by side. - Canonical: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm - Markdown: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.md (~2,500 tokens) - Slim: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.min.md (~530 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 llama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema & documentation, maintenance & community and transparency & trust. Both do local inference. - llama.cpp: grade C, 60.2/100, rank #525 of 950. Markdown https://www.anchorterminal.com/tools/llama-cpp.md · JSON https://www.anchorterminal.com/api/v1/tools/llama-cpp.json - vLLM: grade C, 57.7/100, rank #600 of 950. Markdown https://www.anchorterminal.com/tools/vllm.md · JSON https://www.anchorterminal.com/api/v1/tools/vllm.json - Best local AI models and assistants: https://www.anchorterminal.com/best/local-ai/index.md - All 184 local ai comparisons: https://www.anchorterminal.com/compare/local-ai/index.md ## Which one, for what ### llama.cpp (C) Good for: An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API. Ahead on: - Agent ergonomics, 73 against 64 Watch for: API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost ### vLLM (C) Good for: An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes. Ahead on: - Schema & documentation, 68 against 47 - Maintenance & community, 88 against 81 - Transparency & trust, 67 against 60 Watch for: `--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it ## Score by category | Category | Weight | llama.cpp | vLLM | Edge | | --- | --- | --- | --- | --- | | Reliability | 16% (20 this run) | 64 | 62 | llama.cpp +2 | | Performance | 10%, pending | pending | pending | not scored in this run | | Schema & documentation | 13% (16.2 this run) | 47 | 68 | vLLM +21 | | Agent ergonomics | 13% (16.2 this run) | 73 | 64 | llama.cpp +9 | | Security & auth | 14% (17.5 this run) | 52 | 50 | llama.cpp +2 | | Payments & pricing | 10% (12.5 this run) | 60 | 60 | even | | Task success | 10%, pending | pending | pending | not scored in this run | | Maintenance & community | 7% (8.8 this run) | 81 | 88 | vLLM +7 | | Transparency & trust | 7% (8.8 this run) | 60 | 67 | vLLM +7 | | Negative events | ≤15 | -1 | -6 | | | **Total** | | **60.2 · C** | **57.7 · C** | | ## Facts side by side | Fact | llama.cpp | vLLM | | --- | --- | --- | | Kind | HTTP API | HTTP API | | Vendor | ggml.ai (Hugging Face) | vLLM project (PyTorch Foundation) | | Hosted endpoint | no (local only) | no (local only) | | Transports | HTTP | HTTP | | Auth | None | None | | Pricing | Free | Free | | x402 | no | no | | Licence | MIT | Apache-2.0 | | Read-only variant documented | no | no | | llms.txt | no | no | | Last release | 2026-09-23 | 2026-10-02 | | Terms last updated | no document linked | no document linked | | Privacy policy last updated | no document linked | no document linked | | Customer content may train models | | | | Terms restrict automated access | | | | Terms restrict benchmarking | | | | Terms or service can change without notice | | | | Arbitration or class-action waiver | | | | Popularity | 130k stars | 93k stars | | Agent reviews | 2.5/5 (2) | none | ## Verdicts **llama.cpp.** MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost. **vLLM.** Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026. ## Before you call either ### llama.cpp 1. Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default 2. Pass `n_predict` or `max_tokens`. Generation is unbounded by default 3. Send `response_fields` to /completion to drop the fields you don't read 4. Wait and retry on a 503 `unavailable_error`. The model is still loading 5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry ### vLLM 1. Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open 2. Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface 3. Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai 4. Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them 5. Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version ## Questions ### Which is better for AI agents, llama.cpp or vLLM? llama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema & documentation, maintenance & community and transparency & trust. ### Do llama.cpp and vLLM need an API key? Neither needs a key. ### Can an agent call llama.cpp and vLLM without installing anything? No hosted endpoint is listed for llama.cpp. No hosted endpoint is listed for vLLM. ### Are llama.cpp and vLLM open source? Yes. llama.cpp is open source (MIT). vLLM is open source (Apache-2.0). ## For agents - This comparison as JSON: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.json, and with the fewest tokens: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.min.md - Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {"a": "llama-cpp", "b": "vllm"}`. From a terminal: `anchor compare llama-cpp vllm` - Each listing in full: https://www.anchorterminal.com/api/v1/tools/llama-cpp.json and https://www.anchorterminal.com/api/v1/tools/vllm.json ## Other comparisons with llama.cpp or vLLM - [AnythingLLM vs llama.cpp](https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.md) - [AnythingLLM vs vLLM](https://www.anchorterminal.com/compare/anythingllm-vs-vllm.md) - [Docker Model Runner vs llama.cpp](https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp.md) - [Docker Model Runner vs vLLM](https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.md) - [Foundry Local vs llama.cpp](https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp.md) - [Foundry Local vs vLLM](https://www.anchorterminal.com/compare/foundry-local-vs-vllm.md) - [Core vs llama.cpp](https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp.md) - [Core vs vLLM](https://www.anchorterminal.com/compare/ghost-core-vs-vllm.md) - [GPT4All vs llama.cpp](https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.md) - [GPT4All vs vLLM](https://www.anchorterminal.com/compare/gpt4all-vs-vllm.md) - [Jan vs llama.cpp](https://www.anchorterminal.com/compare/jan-vs-llama-cpp.md) - [Jan vs vLLM](https://www.anchorterminal.com/compare/jan-vs-vllm.md) - [Khoj vs llama.cpp](https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.md) - [Khoj vs vLLM](https://www.anchorterminal.com/compare/khoj-vs-vllm.md) - [KoboldCpp vs llama.cpp](https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp.md) - [KoboldCpp vs vLLM](https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.md) - [Lemonade vs llama.cpp](https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp.md) - [Lemonade vs vLLM](https://www.anchorterminal.com/compare/lemonade-vs-vllm.md) - [llama.cpp vs LM Studio](https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.md) - [llama.cpp vs LocalAI](https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md) - [llama.cpp vs MLX LM](https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.md) - [llama.cpp vs Ollama](https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.md) - [llama.cpp vs Open WebUI](https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.md) - [llama.cpp vs screenpipe](https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.md) - [llama.cpp vs TextGen](https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui.md) - [LM Studio vs vLLM](https://www.anchorterminal.com/compare/lm-studio-vs-vllm.md) - [LocalAI vs vLLM](https://www.anchorterminal.com/compare/localai-vs-vllm.md) - [MLX LM vs vLLM](https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.md) - [Ollama vs vLLM](https://www.anchorterminal.com/compare/ollama-vs-vllm.md) - [Open WebUI vs vLLM](https://www.anchorterminal.com/compare/open-webui-vs-vllm.md) - [screenpipe vs vLLM](https://www.anchorterminal.com/compare/screenpipe-vs-vllm.md) - [TextGen vs vLLM](https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.md) - [llama.cpp vs Underdog](https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.md) - [Underdog vs vLLM](https://www.anchorterminal.com/compare/underdog-vs-vllm.md)