{
  "data": {
    "a": {
      "slug": "llama-cpp",
      "name": "llama.cpp",
      "vendor": "ggml.ai (Hugging Face)",
      "vendorUrl": "https://llama.app",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source C/C++ engine for running GGUF models locally, with a web interface and compatible model APIs.",
      "url": "https://www.anchorterminal.com/tools/llama-cpp",
      "markdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/llama-cpp.json",
      "repo": "https://github.com/ggml-org/llama.cpp",
      "license": "MIT",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "ghcr.io/ggml-org/llama.cpp"
        },
        {
          "registry": "pypi",
          "name": "gguf"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one key or a comma-separated list) or `--api-key-file` (one key a line) turns on a check for every route but /health and the web UI's files, with the key sent as `Authorization: Bearer` or `X-Api-Key`, never in the query string. Keys have no scopes and change only with a restart. TLS is built in with `--ssl-key-file` and `--ssl-cert-file`. The server binds 127.0.0.1:8080 by default, and CORS reflects any Origin with credentials allowed unless built-in tools, MCP servers or `--agent` are on, when it narrows to localhost (https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).",
      "pricing": "free",
      "pricingNotes": "Free under MIT, with no account, key or card. Nothing is sold. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 130200,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "no-telemetry"
      ],
      "lastRelease": "2026-09-23",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 60.2,
        "grade": "C",
        "agentReady": false,
        "rank": 525,
        "ranked": true,
        "rankOf": 950,
        "categoryRank": 6,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 60
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -1,
        "negativeNotes": [
          "2026-03-26. GHSA-j8rj-fmpv-wcxw (CVE-2026-34159, 9.8 at NVD), unauthenticated code execution through a GRAPH_COMPUTE bypass in the RPC backend, the most serious of four advisories published between January and March 2026 (the others a llama-server out-of-bounds write through a negative `n_discard` and two GGUF integer overflows). All were fixed in named builds and published as advisories, SECURITY.md says not to expose the RPC server or llama-server to untrusted networks, and the newest is more than six months old, -1. https://github.com/ggml-org/llama.cpp/security/advisories/GHSA-j8rj-fmpv-wcxw; https://github.com/ggml-org/llama.cpp/security"
        ],
        "verdict": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
        "bestFor": "An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.",
        "strengths": [
          "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads",
          "OpenAI chat completions, responses and embeddings, Anthropic messages, reranking and /v1/systemone from one server",
          "`response_fields`, `json_schema` and `grammar` control the size and shape of output, and errors carry an OpenAI-style type and code",
          "1,005 nightly builds and eight semver releases in 90 days, with 37 workflows running on every push to master",
          "Ten published GitHub advisories with CVEs and fixed builds, and SECURITY.md guidance on untrusted models and inputs"
        ],
        "weaknesses": [
          "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost",
          "No OpenAPI file of its own, and the REST API changelog stops at b4599",
          "Private security disclosure disabled since 1 June 2026, with fixes asked for as public pull requests",
          "Pre-1.0 (0.5.0), and semver releases are bare tags with no notes",
          "No official client library, and `n_predict` defaults to unlimited"
        ],
        "agentNotes": [
          "Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default",
          "Pass `n_predict` or `max_tokens`. Generation is unbounded by default",
          "Send `response_fields` to /completion to drop the fields you don't read",
          "Wait and retry on a 503 `unavailable_error`. The model is still loading",
          "Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 2.5,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 60.2
          }
        ],
        "editorialScores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 66
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "curl -LsSf https://llama.app/install.sh | sh   # or: brew install llama.cpp; winget install llama.cpp\nllama serve -hf ggml-org/Qwen3.5-0.8B-GGUF   # listens on 127.0.0.1:8080",
        "http": "curl --request POST \\\n    --url http://localhost:8080/completion \\\n    --header \"Content-Type: application/json\" \\\n    --data '{\"prompt\": \"Building a website can be done in 10 simple steps:\",\"n_predict\": 128}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/llama-cpp"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "ggml.ai, part of Hugging Face since 2026",
        "domain": "llama.app",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/ggml-org/llama.cpp/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "The repository's About link is llama.app, which says it's by the llama.cpp team and Hugging Face and links no terms, privacy or security page. ggml.ai says the company was acquired by Hugging Face in 2026 and names no address.",
          "The `LICENSE` file reads Copyright (c) 2023-2026 The ggml authors.",
          "llama.app/.well-known/security.txt and llama.app/llms.txt return 404. SECURITY.md points to GitHub private advisories while saying private disclosure is disabled.",
          "There's no shared hosted endpoint. The server runs on the owner's machine."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/llama-cpp.json",
      "live": {
        "slug": "llama-cpp",
        "versions": [
          {
            "registry": "github",
            "name": "ggml-org/llama.cpp",
            "version": "v0.6.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:02:41.971419102Z"
          },
          {
            "registry": "pypi",
            "name": "gguf",
            "version": "0.19.0",
            "released": "2026-05-06",
            "seenAt": "2026-10-09T17:02:41.856683921Z"
          }
        ],
        "githubStars": 130648,
        "pypiWeekly": 1235251,
        "securityTxt": {
          "url": "https://llama.app/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-09T15:39:18.252846717Z"
        },
        "domain": {
          "domain": "llama.app",
          "registered": "2018-07-18",
          "source": "https://pubapi.registry.google/rdap/domain/llama.app",
          "checkedAt": "2026-10-04T13:04:03.05886804Z"
        },
        "updatedAt": "2026-10-09T17:02:41.971419102Z"
      }
    },
    "answer": "llama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema \u0026 documentation, maintenance \u0026 community and transparency \u0026 trust.",
    "b": {
      "slug": "vllm",
      "name": "vLLM",
      "vendor": "vLLM project (PyTorch Foundation)",
      "vendorUrl": "https://vllm.ai",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "vLLM is an open-source inference and serving engine for open-weight language models. `vllm serve` runs an HTTP server with OpenAI-compatible, Anthropic Messages, embedding, reranking and transcription routes on the owner's own GPUs or CPUs.",
      "url": "https://www.anchorterminal.com/tools/vllm",
      "markdownUrl": "https://www.anchorterminal.com/tools/vllm.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/vllm.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/vllm.json",
      "repo": "https://github.com/vllm-project/vllm",
      "license": "Apache-2.0",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "pypi",
          "name": "vllm"
        },
        {
          "registry": "oci",
          "name": "vllm/vllm-openai"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one or several keys) or `VLLM_API_KEY` turns on a Bearer check for paths under `/v1`, `/v2`, `/inference` and `/cohere` only, so `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank` and control routes such as `/pause` stay open. Keys have no scopes and change with a restart. The key is read from the `Authorization` header, never the query string. gRPC has no authentication (https://github.com/vllm-project/vllm/blob/main/docs/usage/security.md).",
      "pricing": "free",
      "pricingNotes": "Free under Apache-2.0, with no account, key or card. Nothing is sold by the project. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-09).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 93444,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-09"
      },
      "docsUrl": "https://docs.vllm.ai/en/stable/",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "speech.stt",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "telemetry-default-on"
      ],
      "lastRelease": "2026-10-02",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 57.7,
        "grade": "C",
        "agentReady": false,
        "rank": 600,
        "ranked": true,
        "rankOf": 950,
        "categoryRank": 8,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 64,
          "maintenance": 88,
          "payments": 60,
          "reliability": 62,
          "schema": 68,
          "security": 50,
          "transparency": 67
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-09"
        },
        "negative": -6,
        "negativeNotes": [
          "2026-06-02. GHSA-94f4-hr76-p5j6 (CVE-2026-48746, 9.1), a crafted Host header bypassed the API key check on the OpenAI routes, fixed in 0.22.0. With GHSA-4r2x-xpjr-7cvv (CVE-2026-22778, 9.8) of 2 February 2026, code execution through video decoding fixed in 0.14.1, these are the two critical advisories of the last 12 months. Both were fixed and published with CVEs, so they decay, -3. https://github.com/vllm-project/vllm/security/advisories/GHSA-94f4-hr76-p5j6; https://github.com/vllm-project/vllm/security/advisories/GHSA-4r2x-xpjr-7cvv",
          "2026-10-06. GHSA-h3rc-6mm3-gc2m (8.1), a request field could select the processor code a server started with `--trust-remote-code` imports, fixed in 0.31.0, one of 50 advisories published since 11 July 2026 (10 high, 36 medium, 4 low), most of them requests that crash or exhaust the engine. All name a fixed version, and eleven were published on 9 October 2026 months after their fixes, -3. https://github.com/vllm-project/vllm/security/advisories/GHSA-h3rc-6mm3-gc2m; https://github.com/vllm-project/vllm/security/advisories"
        ],
        "verdict": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.",
        "bestFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "strengths": [
          "OpenAI chat, completions, responses and embeddings, Anthropic `/v1/messages`, Cohere embed and rerank, transcription and `/v1/systemone` from one server",
          "Apache-2.0, with a written three-stage deprecation policy and release notes that carry a breaking changes section",
          "Eight stable releases between 12 July and 2 October 2026, and v0.31.0 lists 717 commits from 307 contributors",
          "A 650-line security guide names every route the API key does and does not protect, and the limits of multi-tenant use",
          "Usage statistics are documented field by field, with `VLLM_NO_USAGE_STATS`, `DO_NOT_TRACK` or a file as opt-outs"
        ],
        "weaknesses": [
          "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it",
          "No key by default, the server binds every interface when `--host` is unset, and CORS allows any origin",
          "At least 81 GitHub security advisories in 12 months, two critical, most of them remote crashes or resource exhaustion",
          "Pre-1.0 (0.31.0), with breaking changes in each fortnightly release and compatibility kept for a limited number of minor versions",
          "Usage statistics are sent to stats.vllm.ai by default, and no privacy policy or retention period for them was found"
        ],
        "agentNotes": [
          "Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open",
          "Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface",
          "Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai",
          "Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them",
          "Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 57.7
          }
        ],
        "editorialScores": {
          "ergonomics": 64,
          "maintenance": 88,
          "payments": 60,
          "reliability": 62,
          "schema": 68,
          "security": 50,
          "transparency": 80
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "uv pip install vllm --torch-backend=auto\nvllm serve Qwen/Qwen2.5-1.5B-Instruct   # listens on port 8000",
        "http": "curl http://localhost:8000/v1/chat/completions \\\n    -H \"Content-Type: application/json\" \\\n    -d '{\n        \"model\": \"Qwen/Qwen2.5-1.5B-Instruct\",\n        \"messages\": [\n            {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n            {\"role\": \"user\", \"content\": \"Who won the world series in 2020?\"}\n        ]\n    }'",
        "claudeCode": "ANTHROPIC_BASE_URL=http://localhost:8000 \\\nANTHROPIC_API_KEY=dummy \\\nANTHROPIC_AUTH_TOKEN=dummy \\\nANTHROPIC_DEFAULT_OPUS_MODEL=my-model \\\nANTHROPIC_DEFAULT_SONNET_MODEL=my-model \\\nANTHROPIC_DEFAULT_HAIKU_MODEL=my-model \\\nclaude"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/vllm"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "The Linux Foundation (vLLM is a PyTorch Foundation project)",
        "domain": "vllm.ai",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/vllm-project/vllm/releases",
        "securityTxt": "none",
        "checked": "2026-10-09",
        "notes": [
          "vllm.ai links no terms and no privacy policy, and its footer reads © 2026 vLLM. The project publishes none for the software or for stats.vllm.ai, so the Apache-2.0 licence stands in for terms.",
          "pytorch.org/projects/vllm/ lists vLLM among PyTorch Foundation projects and says UC Berkeley contributed it to the Linux Foundation in July 2024. The Linux Foundation's policies are linked from that page and are not specific to vLLM.",
          "vllm.ai/.well-known/security.txt and docs.vllm.ai/.well-known/security.txt return 404. SECURITY.md asks for private reports through GitHub.",
          "There's no shared hosted endpoint. The server runs on the owner's hardware. The software posts usage statistics to stats.vllm.ai unless turned off."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/vllm.json",
      "live": {
        "slug": "vllm",
        "versions": [
          {
            "registry": "github",
            "name": "vllm-project/vllm",
            "version": "v0.31.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:27:30.444059802Z"
          },
          {
            "registry": "pypi",
            "name": "vllm",
            "version": "0.31.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:27:30.259454009Z"
          }
        ],
        "githubStars": 93457,
        "pypiWeekly": 444466,
        "updatedAt": "2026-10-09T17:27:30.444059802Z"
      }
    },
    "facts": [
      {
        "a": "HTTP API",
        "b": "HTTP API",
        "name": "Kind"
      },
      {
        "a": "ggml.ai (Hugging Face)",
        "b": "vLLM project (PyTorch Foundation)",
        "name": "Vendor"
      },
      {
        "a": "no (local only)",
        "b": "no (local only)",
        "name": "Hosted endpoint"
      },
      {
        "a": "HTTP",
        "b": "HTTP",
        "name": "Transports"
      },
      {
        "a": "None",
        "b": "None",
        "name": "Auth"
      },
      {
        "a": "Free",
        "b": "Free",
        "name": "Pricing"
      },
      {
        "a": "no",
        "b": "no",
        "name": "x402"
      },
      {
        "a": "MIT",
        "b": "Apache-2.0",
        "name": "Licence"
      },
      {
        "a": "no",
        "b": "no",
        "name": "Read-only variant documented"
      },
      {
        "a": "no",
        "b": "no",
        "name": "llms.txt"
      },
      {
        "a": "2026-09-23",
        "b": "2026-10-02",
        "name": "Last release"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Terms last updated"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Privacy policy last updated"
      },
      {
        "a": "",
        "b": "",
        "name": "Customer content may train models"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict automated access"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict benchmarking"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms or service can change without notice"
      },
      {
        "a": "",
        "b": "",
        "name": "Arbitration or class-action waiver"
      },
      {
        "a": "130k stars",
        "b": "93k stars",
        "name": "Popularity"
      },
      {
        "a": "2.5/5 (2)",
        "b": "none",
        "name": "Agent reviews"
      }
    ],
    "faq": [
      {
        "answer": "llama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema \u0026 documentation, maintenance \u0026 community and transparency \u0026 trust.",
        "question": "Which is better for AI agents, llama.cpp or vLLM?"
      },
      {
        "answer": "Neither needs a key.",
        "question": "Do llama.cpp and vLLM need an API key?"
      },
      {
        "answer": "No hosted endpoint is listed for llama.cpp. No hosted endpoint is listed for vLLM.",
        "question": "Can an agent call llama.cpp and vLLM without installing anything?"
      },
      {
        "answer": "Yes. llama.cpp is open source (MIT). vLLM is open source (Apache-2.0).",
        "question": "Are llama.cpp and vLLM open source?"
      }
    ],
    "goodFor": [
      {
        "aheadOn": [
          "Agent ergonomics, 73 against 64"
        ],
        "also": null,
        "goodFor": "An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.",
        "slug": "llama-cpp",
        "watchFor": "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost"
      },
      {
        "aheadOn": [
          "Schema \u0026 documentation, 68 against 47",
          "Maintenance \u0026 community, 88 against 81",
          "Transparency \u0026 trust, 67 against 60"
        ],
        "also": null,
        "goodFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "slug": "vllm",
        "watchFor": "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it"
      }
    ],
    "job": {
      "capability": "inference.local",
      "name": "Local inference"
    },
    "others": [
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.json",
        "title": "AnythingLLM vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-vllm.json",
        "title": "AnythingLLM vs vLLM",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp.json",
        "title": "Docker Model Runner vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.json",
        "title": "Docker Model Runner vs vLLM",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp.json",
        "title": "Foundry Local vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-vllm.json",
        "title": "Foundry Local vs vLLM",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp.json",
        "title": "Core vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-vllm.json",
        "title": "Core vs vLLM",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.json",
        "title": "GPT4All vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-vllm.json",
        "title": "GPT4All vs vLLM",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-llama-cpp.json",
        "title": "Jan vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/jan-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-vllm.json",
        "title": "Jan vs vLLM",
        "url": "https://www.anchorterminal.com/compare/jan-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.json",
        "title": "Khoj vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-vllm.json",
        "title": "Khoj vs vLLM",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp.json",
        "title": "KoboldCpp vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.json",
        "title": "KoboldCpp vs vLLM",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp.json",
        "title": "Lemonade vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-vllm.json",
        "title": "Lemonade vs vLLM",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.json",
        "title": "llama.cpp vs LM Studio",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai.json",
        "title": "llama.cpp vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.json",
        "title": "llama.cpp vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.json",
        "title": "llama.cpp vs Ollama",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.json",
        "title": "llama.cpp vs Open WebUI",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.json",
        "title": "llama.cpp vs screenpipe",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui.json",
        "title": "llama.cpp vs TextGen",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lm-studio-vs-vllm.json",
        "title": "LM Studio vs vLLM",
        "url": "https://www.anchorterminal.com/compare/lm-studio-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-vllm.json",
        "title": "LocalAI vs vLLM",
        "url": "https://www.anchorterminal.com/compare/localai-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.json",
        "title": "MLX LM vs vLLM",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ollama-vs-vllm.json",
        "title": "Ollama vs vLLM",
        "url": "https://www.anchorterminal.com/compare/ollama-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/open-webui-vs-vllm.json",
        "title": "Open WebUI vs vLLM",
        "url": "https://www.anchorterminal.com/compare/open-webui-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/screenpipe-vs-vllm.json",
        "title": "screenpipe vs vLLM",
        "url": "https://www.anchorterminal.com/compare/screenpipe-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.json",
        "title": "TextGen vs vLLM",
        "url": "https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.json",
        "title": "llama.cpp vs Underdog",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-underdog"
      },
      {
        "json": "https://www.anchorterminal.com/compare/underdog-vs-vllm.json",
        "title": "Underdog vs vLLM",
        "url": "https://www.anchorterminal.com/compare/underdog-vs-vllm"
      }
    ],
    "scores": [
      {
        "by": 2,
        "edge": "llama-cpp",
        "key": "reliability",
        "llama-cpp": 64,
        "name": "Reliability",
        "vllm": 62,
        "weight": 16
      },
      {
        "key": "performance",
        "name": "Performance",
        "pending": true,
        "weight": 10
      },
      {
        "by": 21,
        "edge": "vllm",
        "key": "schema",
        "llama-cpp": 47,
        "name": "Schema \u0026 documentation",
        "vllm": 68,
        "weight": 13
      },
      {
        "by": 9,
        "edge": "llama-cpp",
        "key": "ergonomics",
        "llama-cpp": 73,
        "name": "Agent ergonomics",
        "vllm": 64,
        "weight": 13
      },
      {
        "by": 2,
        "edge": "llama-cpp",
        "key": "security",
        "llama-cpp": 52,
        "name": "Security \u0026 auth",
        "vllm": 50,
        "weight": 14
      },
      {
        "by": 0,
        "edge": "",
        "key": "payments",
        "llama-cpp": 60,
        "name": "Payments \u0026 pricing",
        "vllm": 60,
        "weight": 10
      },
      {
        "key": "tasks",
        "name": "Task success",
        "pending": true,
        "weight": 10
      },
      {
        "by": 7,
        "edge": "vllm",
        "key": "maintenance",
        "llama-cpp": 81,
        "name": "Maintenance \u0026 community",
        "vllm": 88,
        "weight": 7
      },
      {
        "by": 7,
        "edge": "vllm",
        "key": "transparency",
        "llama-cpp": 60,
        "name": "Transparency \u0026 trust",
        "vllm": 67,
        "weight": 7
      }
    ],
    "summary": "llama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema \u0026 documentation, maintenance \u0026 community and transparency \u0026 trust. Both do local inference.",
    "verdicts": {
      "llama-cpp": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
      "vllm": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026."
    }
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm",
    "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.md",
    "slim": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.min.md"
  },
  "markdown": "llama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema \u0026 documentation, maintenance \u0026 community and transparency \u0026 trust. Both do local inference.\n\n- llama.cpp: grade C, 60.2/100, rank #525 of 950. Markdown https://www.anchorterminal.com/tools/llama-cpp.md · JSON https://www.anchorterminal.com/api/v1/tools/llama-cpp.json\n- vLLM: grade C, 57.7/100, rank #600 of 950. Markdown https://www.anchorterminal.com/tools/vllm.md · JSON https://www.anchorterminal.com/api/v1/tools/vllm.json\n- Best local AI models and assistants: https://www.anchorterminal.com/best/local-ai/index.md\n- All 184 local ai comparisons: https://www.anchorterminal.com/compare/local-ai/index.md\n\n## Which one, for what\n\n### llama.cpp (C)\n\nGood for: An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.\n\nAhead on:\n- Agent ergonomics, 73 against 64\n\nWatch for: API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost\n\n### vLLM (C)\n\nGood for: An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.\n\nAhead on:\n- Schema \u0026 documentation, 68 against 47\n- Maintenance \u0026 community, 88 against 81\n- Transparency \u0026 trust, 67 against 60\n\nWatch for: `--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it\n\n\n## Score by category\n\n| Category | Weight | llama.cpp | vLLM | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 64 | 62 | llama.cpp +2 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 47 | 68 | vLLM +21 |\n| Agent ergonomics | 13% (16.2 this run) | 73 | 64 | llama.cpp +9 |\n| Security \u0026 auth | 14% (17.5 this run) | 52 | 50 | llama.cpp +2 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 60 | 60 | even |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 81 | 88 | vLLM +7 |\n| Transparency \u0026 trust | 7% (8.8 this run) | 60 | 67 | vLLM +7 |\n| Negative events | ≤15 | -1 | -6 | |\n| **Total** | | **60.2 · C** | **57.7 · C** | |\n\n## Facts side by side\n\n| Fact | llama.cpp | vLLM |\n| --- | --- | --- |\n| Kind | HTTP API | HTTP API |\n| Vendor | ggml.ai (Hugging Face) | vLLM project (PyTorch Foundation) |\n| Hosted endpoint | no (local only) | no (local only) |\n| Transports | HTTP | HTTP |\n| Auth | None | None |\n| Pricing | Free | Free |\n| x402 | no | no |\n| Licence | MIT | Apache-2.0 |\n| Read-only variant documented | no | no |\n| llms.txt | no | no |\n| Last release | 2026-09-23 | 2026-10-02 |\n| Terms last updated | no document linked | no document linked |\n| Privacy policy last updated | no document linked | no document linked |\n| Customer content may train models |  |  |\n| Terms restrict automated access |  |  |\n| Terms restrict benchmarking |  |  |\n| Terms or service can change without notice |  |  |\n| Arbitration or class-action waiver |  |  |\n| Popularity | 130k stars | 93k stars |\n| Agent reviews | 2.5/5 (2) | none |\n\n## Verdicts\n\n**llama.cpp.** MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.\n\n**vLLM.** Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.\n\n## Before you call either\n\n### llama.cpp\n\n1. Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default\n2. Pass `n_predict` or `max_tokens`. Generation is unbounded by default\n3. Send `response_fields` to /completion to drop the fields you don't read\n4. Wait and retry on a 503 `unavailable_error`. The model is still loading\n5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry\n\n### vLLM\n\n1. Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open\n2. Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface\n3. Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai\n4. Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them\n5. Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version\n\n## Questions\n\n### Which is better for AI agents, llama.cpp or vLLM?\n\nllama.cpp scores 60.2 (C) on agent readiness against vLLM's 57.7 (C), and leads in 3 of 7 scored categories. vLLM leads on schema \u0026 documentation, maintenance \u0026 community and transparency \u0026 trust.\n\n### Do llama.cpp and vLLM need an API key?\n\nNeither needs a key.\n\n### Can an agent call llama.cpp and vLLM without installing anything?\n\nNo hosted endpoint is listed for llama.cpp. No hosted endpoint is listed for vLLM.\n\n### Are llama.cpp and vLLM open source?\n\nYes. llama.cpp is open source (MIT). vLLM is open source (Apache-2.0).\n\n\n## For agents\n\n- This comparison as JSON: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.json, and with the fewest tokens: https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.min.md\n- Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {\"a\": \"llama-cpp\", \"b\": \"vllm\"}`. From a terminal: `anchor compare llama-cpp vllm`\n- Each listing in full: https://www.anchorterminal.com/api/v1/tools/llama-cpp.json and https://www.anchorterminal.com/api/v1/tools/vllm.json\n\n## Other comparisons with llama.cpp or vLLM\n\n- [AnythingLLM vs llama.cpp](https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.md)\n- [AnythingLLM vs vLLM](https://www.anchorterminal.com/compare/anythingllm-vs-vllm.md)\n- [Docker Model Runner vs llama.cpp](https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp.md)\n- [Docker Model Runner vs vLLM](https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.md)\n- [Foundry Local vs llama.cpp](https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp.md)\n- [Foundry Local vs vLLM](https://www.anchorterminal.com/compare/foundry-local-vs-vllm.md)\n- [Core vs llama.cpp](https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp.md)\n- [Core vs vLLM](https://www.anchorterminal.com/compare/ghost-core-vs-vllm.md)\n- [GPT4All vs llama.cpp](https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.md)\n- [GPT4All vs vLLM](https://www.anchorterminal.com/compare/gpt4all-vs-vllm.md)\n- [Jan vs llama.cpp](https://www.anchorterminal.com/compare/jan-vs-llama-cpp.md)\n- [Jan vs vLLM](https://www.anchorterminal.com/compare/jan-vs-vllm.md)\n- [Khoj vs llama.cpp](https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.md)\n- [Khoj vs vLLM](https://www.anchorterminal.com/compare/khoj-vs-vllm.md)\n- [KoboldCpp vs llama.cpp](https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp.md)\n- [KoboldCpp vs vLLM](https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.md)\n- [Lemonade vs llama.cpp](https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp.md)\n- [Lemonade vs vLLM](https://www.anchorterminal.com/compare/lemonade-vs-vllm.md)\n- [llama.cpp vs LM Studio](https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.md)\n- [llama.cpp vs LocalAI](https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md)\n- [llama.cpp vs MLX LM](https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.md)\n- [llama.cpp vs Ollama](https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.md)\n- [llama.cpp vs Open WebUI](https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.md)\n- [llama.cpp vs screenpipe](https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.md)\n- [llama.cpp vs TextGen](https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui.md)\n- [LM Studio vs vLLM](https://www.anchorterminal.com/compare/lm-studio-vs-vllm.md)\n- [LocalAI vs vLLM](https://www.anchorterminal.com/compare/localai-vs-vllm.md)\n- [MLX LM vs vLLM](https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.md)\n- [Ollama vs vLLM](https://www.anchorterminal.com/compare/ollama-vs-vllm.md)\n- [Open WebUI vs vLLM](https://www.anchorterminal.com/compare/open-webui-vs-vllm.md)\n- [screenpipe vs vLLM](https://www.anchorterminal.com/compare/screenpipe-vs-vllm.md)\n- [TextGen vs vLLM](https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.md)\n- [llama.cpp vs Underdog](https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.md)\n- [Underdog vs vLLM](https://www.anchorterminal.com/compare/underdog-vs-vllm.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-10",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "llama.cpp vs vLLM",
        "url": ""
      }
    ],
    "description": "llama.cpp scores 60.2 (C) to vLLM's 57.7 (C) for local inference. Prices, MCP, x402, uptime and agent notes side by side.",
    "facts": [
      "llama.cpp C 60.2",
      "vLLM C 57.7",
      "scores"
    ],
    "h1": "llama.cpp vs vLLM",
    "image": "https://www.anchorterminal.com/assets/og/compare-llama-cpp-vs-vllm.png",
    "path": "/compare/llama-cpp-vs-vllm",
    "published": "2026-10-01",
    "section": "tools",
    "title": "llama.cpp vs vLLM for AI agents in 2026: scores and prices",
    "toc": null,
    "updated": "2026-10-09",
    "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm"
  },
  "tokens": {
    "markdown": 2500,
    "slim": 530
  },
  "version": 1
}
