{
  "data": {
    "a": {
      "slug": "localai",
      "name": "LocalAI",
      "vendor": "Ettore Di Giacinto and the LocalAI team",
      "vendorUrl": "https://localai.io",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source engine in Go, MIT licensed, that runs models on the owner's hardware behind OpenAI-, Anthropic-, Ollama- and ElevenLabs-compatible APIs on port 8080.",
      "url": "https://www.anchorterminal.com/tools/localai",
      "markdownUrl": "https://www.anchorterminal.com/tools/localai.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/localai.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/localai.json",
      "repo": "https://github.com/mudler/LocalAI",
      "license": "MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence",
      "transports": [
        "http",
        "stdio"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "docker.io/localai/localai"
        }
      ],
      "auth": "mixed",
      "authNotes": "Off by default. With no keys and no user accounts configured, every request is accepted, and the server refuses to start on a public address in that state unless `--allow-insecure-public-bind` is set. `LOCALAI_API_KEY` sets shared keys with full admin rights. `LOCALAI_AUTH=true` turns on user accounts (local, GitHub OAuth or OIDC), and each user creates revocable keys stored as HMAC-SHA256, with an optional expiry in the source, carrying the user's role (admin or user) and per-model and per-feature permissions. Keys go in `Authorization: Bearer`, `x-api-key`, `xi-api-key` or a `token` cookie.",
      "pricing": "free",
      "pricingNotes": "Free and MIT with nothing to buy. You run it on your own hardware. The project takes sponsorship through GitHub Sponsors.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": 42,
      "popularity": {
        "githubStars": 47800,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://localai.io/basics/getting_started/",
      "openapi": "https://raw.githubusercontent.com/mudler/LocalAI/master/swagger/swagger.json",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "agent.mcp-client",
        "embed.text",
        "rerank",
        "speech.stt",
        "speech.tts",
        "voice.speech-to-speech",
        "image.generate",
        "video.generate",
        "guard.pii",
        "finetune.sft",
        "db.vector"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "openapi",
        "mcp",
        "go",
        "docker",
        "streaming",
        "open-weights",
        "no-telemetry"
      ],
      "lastRelease": "2026-10-02",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 68,
        "grade": "B",
        "agentReady": false,
        "rank": 232,
        "ranked": true,
        "rankOf": 950,
        "categoryRank": 1,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 71,
          "maintenance": 80,
          "payments": 60,
          "reliability": 84,
          "schema": 81,
          "security": 62,
          "transparency": 47
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -3,
        "negativeNotes": [
          "2026-07-07. CVE-2026-59707 (8.6 at NVD under CVSS 3.1, published by VulnCheck), an unauthenticated server-side request forgery through POST /models/apply in v4.3.1 and earlier, reported in issue #10665. The code now refuses private, loopback and metadata addresses in gallery config fetches, with a comment citing the issue, from v4.8.0 at the latest. The project published no GitHub advisory, the fix commit NVD and VulnCheck name (f9b968e) is an unrelated docs change, and SECURITY.md still lists 3.x as the supported series. Fixed, documented only by a third party, -3. https://nvd.nist.gov/vuln/detail/CVE-2026-59707; https://github.com/mudler/LocalAI/issues/10665"
        ],
        "verdict": "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.",
        "bestFor": "An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator.",
        "strengths": [
          "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app",
          "OpenAI, Anthropic, Open Responses, Ollama and ElevenLabs-compatible endpoints, with a Swagger 2.0 file of 133 operations served by every instance",
          "429 and 503 responses carry Retry-After, and errors come in the calling client's own envelope",
          "Optional user accounts with hashed, revocable keys, per-model and per-feature permissions and per-user quotas",
          "v4.11.0 on 2 October 2026, ten releases in 90 days, and the Tests workflow passing on master"
        ],
        "weaknesses": [
          "No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin",
          "CVE-2026-59707, an unauthenticated SSRF in v4.3.1 and earlier, published by VulnCheck in July 2026 with no advisory from the project",
          "SECURITY.md still names 3.x as the supported series, and there's no security.txt or privacy policy",
          "The MCP admin server registers 42 tools against the 19 its docs list, with no annotations, and its writes are held back only by a prompt",
          "No breaking-change section in the release notes, and the unsigned macOS DMG needs its quarantine flag removed by hand"
        ],
        "agentNotes": [
          "Send `Authorization: Bearer \u003ckey\u003e` when the operator has set keys. A 401 means the instance has auth on",
          "Read /.well-known/localai.json and /api/instructions first. Both answer without a key and list what this instance can do",
          "Back off on 429 and 503 for the Retry-After seconds. A 503 can mean the model is still loading",
          "Start `local-ai mcp-server` with `--read-only` unless the task is to install or delete models",
          "Take model names from /v1/models. Each instance names its own"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 3,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "B",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 68
          }
        ],
        "editorialScores": {
          "ergonomics": 71,
          "maintenance": 80,
          "payments": 60,
          "reliability": 84,
          "schema": 81,
          "security": 62,
          "transparency": 67
        },
        "provenanceScore": 27
      },
      "connect": {
        "install": "docker run -ti --name local-ai -p 8080:8080 localai/localai:latest",
        "http": "curl http://localhost:8080/v1/chat/completions -H \"Content-Type: application/json\" -d '{\n  \"model\": \"qwen3-4b\",\n  \"messages\": [{\"role\": \"user\", \"content\": \"Hello!\"}]\n}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/localai"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "",
        "domain": "localai.io",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/mudler/LocalAI/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "No company is named. The `LICENSE` copyright line reads Ettore Di Giacinto, and the README names him as project lead with Richard Palethorpe as maintainer.",
          "We found no terms or privacy page in the docs site's source, and localai.io/.well-known/security.txt and localai.io/llms.txt return 404.",
          "There's no hosted endpoint. Each instance answers on the operator's own host, by default port 8080."
        ],
        "score": 27
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/localai.json",
      "live": {
        "slug": "localai",
        "versions": [
          {
            "registry": "github",
            "name": "mudler/LocalAI",
            "version": "v4.11.0",
            "released": "2026-10-02",
            "seenAt": "2026-10-09T17:02:59.531252345Z"
          }
        ],
        "githubStars": 49453,
        "securityTxt": {
          "url": "https://localai.io/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-09T15:39:54.911008842Z"
        },
        "domain": {
          "domain": "localai.io",
          "checkedAt": "2026-10-04T13:07:02.946116654Z"
        },
        "updatedAt": "2026-10-09T17:02:59.531252345Z"
      }
    },
    "answer": "LocalAI scores 68 (B) on agent readiness against vLLM's 57.7 (C), and leads in 4 of 7 scored categories. vLLM leads on maintenance \u0026 community and transparency \u0026 trust.",
    "b": {
      "slug": "vllm",
      "name": "vLLM",
      "vendor": "vLLM project (PyTorch Foundation)",
      "vendorUrl": "https://vllm.ai",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "vLLM is an open-source inference and serving engine for open-weight language models. `vllm serve` runs an HTTP server with OpenAI-compatible, Anthropic Messages, embedding, reranking and transcription routes on the owner's own GPUs or CPUs.",
      "url": "https://www.anchorterminal.com/tools/vllm",
      "markdownUrl": "https://www.anchorterminal.com/tools/vllm.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/vllm.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/vllm.json",
      "repo": "https://github.com/vllm-project/vllm",
      "license": "Apache-2.0",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "pypi",
          "name": "vllm"
        },
        {
          "registry": "oci",
          "name": "vllm/vllm-openai"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one or several keys) or `VLLM_API_KEY` turns on a Bearer check for paths under `/v1`, `/v2`, `/inference` and `/cohere` only, so `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank` and control routes such as `/pause` stay open. Keys have no scopes and change with a restart. The key is read from the `Authorization` header, never the query string. gRPC has no authentication (https://github.com/vllm-project/vllm/blob/main/docs/usage/security.md).",
      "pricing": "free",
      "pricingNotes": "Free under Apache-2.0, with no account, key or card. Nothing is sold by the project. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-09).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 93444,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-09"
      },
      "docsUrl": "https://docs.vllm.ai/en/stable/",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "speech.stt",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "telemetry-default-on"
      ],
      "lastRelease": "2026-10-02",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 57.7,
        "grade": "C",
        "agentReady": false,
        "rank": 600,
        "ranked": true,
        "rankOf": 950,
        "categoryRank": 8,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 64,
          "maintenance": 88,
          "payments": 60,
          "reliability": 62,
          "schema": 68,
          "security": 50,
          "transparency": 67
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-09"
        },
        "negative": -6,
        "negativeNotes": [
          "2026-06-02. GHSA-94f4-hr76-p5j6 (CVE-2026-48746, 9.1), a crafted Host header bypassed the API key check on the OpenAI routes, fixed in 0.22.0. With GHSA-4r2x-xpjr-7cvv (CVE-2026-22778, 9.8) of 2 February 2026, code execution through video decoding fixed in 0.14.1, these are the two critical advisories of the last 12 months. Both were fixed and published with CVEs, so they decay, -3. https://github.com/vllm-project/vllm/security/advisories/GHSA-94f4-hr76-p5j6; https://github.com/vllm-project/vllm/security/advisories/GHSA-4r2x-xpjr-7cvv",
          "2026-10-06. GHSA-h3rc-6mm3-gc2m (8.1), a request field could select the processor code a server started with `--trust-remote-code` imports, fixed in 0.31.0, one of 50 advisories published since 11 July 2026 (10 high, 36 medium, 4 low), most of them requests that crash or exhaust the engine. All name a fixed version, and eleven were published on 9 October 2026 months after their fixes, -3. https://github.com/vllm-project/vllm/security/advisories/GHSA-h3rc-6mm3-gc2m; https://github.com/vllm-project/vllm/security/advisories"
        ],
        "verdict": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.",
        "bestFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "strengths": [
          "OpenAI chat, completions, responses and embeddings, Anthropic `/v1/messages`, Cohere embed and rerank, transcription and `/v1/systemone` from one server",
          "Apache-2.0, with a written three-stage deprecation policy and release notes that carry a breaking changes section",
          "Eight stable releases between 12 July and 2 October 2026, and v0.31.0 lists 717 commits from 307 contributors",
          "A 650-line security guide names every route the API key does and does not protect, and the limits of multi-tenant use",
          "Usage statistics are documented field by field, with `VLLM_NO_USAGE_STATS`, `DO_NOT_TRACK` or a file as opt-outs"
        ],
        "weaknesses": [
          "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it",
          "No key by default, the server binds every interface when `--host` is unset, and CORS allows any origin",
          "At least 81 GitHub security advisories in 12 months, two critical, most of them remote crashes or resource exhaustion",
          "Pre-1.0 (0.31.0), with breaking changes in each fortnightly release and compatibility kept for a limited number of minor versions",
          "Usage statistics are sent to stats.vllm.ai by default, and no privacy policy or retention period for them was found"
        ],
        "agentNotes": [
          "Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open",
          "Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface",
          "Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai",
          "Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them",
          "Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 57.7
          }
        ],
        "editorialScores": {
          "ergonomics": 64,
          "maintenance": 88,
          "payments": 60,
          "reliability": 62,
          "schema": 68,
          "security": 50,
          "transparency": 80
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "uv pip install vllm --torch-backend=auto\nvllm serve Qwen/Qwen2.5-1.5B-Instruct   # listens on port 8000",
        "http": "curl http://localhost:8000/v1/chat/completions \\\n    -H \"Content-Type: application/json\" \\\n    -d '{\n        \"model\": \"Qwen/Qwen2.5-1.5B-Instruct\",\n        \"messages\": [\n            {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n            {\"role\": \"user\", \"content\": \"Who won the world series in 2020?\"}\n        ]\n    }'",
        "claudeCode": "ANTHROPIC_BASE_URL=http://localhost:8000 \\\nANTHROPIC_API_KEY=dummy \\\nANTHROPIC_AUTH_TOKEN=dummy \\\nANTHROPIC_DEFAULT_OPUS_MODEL=my-model \\\nANTHROPIC_DEFAULT_SONNET_MODEL=my-model \\\nANTHROPIC_DEFAULT_HAIKU_MODEL=my-model \\\nclaude"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/vllm"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "The Linux Foundation (vLLM is a PyTorch Foundation project)",
        "domain": "vllm.ai",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/vllm-project/vllm/releases",
        "securityTxt": "none",
        "checked": "2026-10-09",
        "notes": [
          "vllm.ai links no terms and no privacy policy, and its footer reads © 2026 vLLM. The project publishes none for the software or for stats.vllm.ai, so the Apache-2.0 licence stands in for terms.",
          "pytorch.org/projects/vllm/ lists vLLM among PyTorch Foundation projects and says UC Berkeley contributed it to the Linux Foundation in July 2024. The Linux Foundation's policies are linked from that page and are not specific to vLLM.",
          "vllm.ai/.well-known/security.txt and docs.vllm.ai/.well-known/security.txt return 404. SECURITY.md asks for private reports through GitHub.",
          "There's no shared hosted endpoint. The server runs on the owner's hardware. The software posts usage statistics to stats.vllm.ai unless turned off."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/vllm.json",
      "live": {
        "slug": "vllm",
        "versions": [
          {
            "registry": "github",
            "name": "vllm-project/vllm",
            "version": "v0.31.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:27:30.444059802Z"
          },
          {
            "registry": "pypi",
            "name": "vllm",
            "version": "0.31.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:27:30.259454009Z"
          }
        ],
        "githubStars": 93457,
        "pypiWeekly": 444466,
        "updatedAt": "2026-10-09T17:27:30.444059802Z"
      }
    },
    "facts": [
      {
        "a": "HTTP API",
        "b": "HTTP API",
        "name": "Kind"
      },
      {
        "a": "Ettore Di Giacinto and the LocalAI team",
        "b": "vLLM project (PyTorch Foundation)",
        "name": "Vendor"
      },
      {
        "a": "no (local only)",
        "b": "no (local only)",
        "name": "Hosted endpoint"
      },
      {
        "a": "HTTP, stdio",
        "b": "HTTP",
        "name": "Transports"
      },
      {
        "a": "OAuth or key",
        "b": "None",
        "name": "Auth"
      },
      {
        "a": "Free",
        "b": "Free",
        "name": "Pricing"
      },
      {
        "a": "no",
        "b": "no",
        "name": "x402"
      },
      {
        "a": "MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence",
        "b": "Apache-2.0",
        "name": "Licence"
      },
      {
        "a": "42",
        "b": "none",
        "name": "Tools exposed"
      },
      {
        "a": "yes",
        "b": "no",
        "name": "Read-only variant documented"
      },
      {
        "a": "no",
        "b": "no",
        "name": "llms.txt"
      },
      {
        "a": "2026-10-02",
        "b": "2026-10-02",
        "name": "Last release"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Terms last updated"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Privacy policy last updated"
      },
      {
        "a": "",
        "b": "",
        "name": "Customer content may train models"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict automated access"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict benchmarking"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms or service can change without notice"
      },
      {
        "a": "",
        "b": "",
        "name": "Arbitration or class-action waiver"
      },
      {
        "a": "48k stars",
        "b": "93k stars",
        "name": "Popularity"
      },
      {
        "a": "3/5 (2)",
        "b": "none",
        "name": "Agent reviews"
      }
    ],
    "faq": [
      {
        "answer": "LocalAI scores 68 (B) on agent readiness against vLLM's 57.7 (C), and leads in 4 of 7 scored categories. vLLM leads on maintenance \u0026 community and transparency \u0026 trust.",
        "question": "Which is better for AI agents, LocalAI or vLLM?"
      },
      {
        "answer": "LocalAI takes an API key or an OAuth sign-in. vLLM needs no key.",
        "question": "Do LocalAI and vLLM need an API key?"
      },
      {
        "answer": "LocalAI runs on your own machine, with no hosted endpoint listed. No hosted endpoint is listed for vLLM.",
        "question": "Can an agent call LocalAI and vLLM without installing anything?"
      },
      {
        "answer": "Yes. LocalAI is open source (MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence). vLLM is open source (Apache-2.0).",
        "question": "Are LocalAI and vLLM open source?"
      }
    ],
    "goodFor": [
      {
        "aheadOn": [
          "Reliability, 84 against 62",
          "Schema \u0026 documentation, 81 against 68",
          "Agent ergonomics, 71 against 64",
          "Security \u0026 auth, 62 against 50"
        ],
        "also": [
          "Runs on your own machine"
        ],
        "goodFor": "An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator.",
        "slug": "localai",
        "watchFor": "No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin"
      },
      {
        "aheadOn": [
          "Maintenance \u0026 community, 88 against 80",
          "Transparency \u0026 trust, 67 against 47"
        ],
        "also": [
          "No key needed to call it"
        ],
        "goodFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "slug": "vllm",
        "watchFor": "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it"
      }
    ],
    "job": {
      "capability": "inference.local",
      "name": "Local inference"
    },
    "others": [
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-localai.json",
        "title": "AnythingLLM vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-vllm.json",
        "title": "AnythingLLM vs vLLM",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-localai.json",
        "title": "Docker Model Runner vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.json",
        "title": "Docker Model Runner vs vLLM",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-localai.json",
        "title": "Foundry Local vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-vllm.json",
        "title": "Foundry Local vs vLLM",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-localai.json",
        "title": "Core vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-vllm.json",
        "title": "Core vs vLLM",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-localai.json",
        "title": "GPT4All vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-vllm.json",
        "title": "GPT4All vs vLLM",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-localai.json",
        "title": "Jan vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/jan-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-vllm.json",
        "title": "Jan vs vLLM",
        "url": "https://www.anchorterminal.com/compare/jan-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-localai.json",
        "title": "Khoj vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-vllm.json",
        "title": "Khoj vs vLLM",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-localai.json",
        "title": "KoboldCpp vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.json",
        "title": "KoboldCpp vs vLLM",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-localai.json",
        "title": "Lemonade vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-vllm.json",
        "title": "Lemonade vs vLLM",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai.json",
        "title": "llama.cpp vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.json",
        "title": "llama.cpp vs vLLM",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lm-studio-vs-localai.json",
        "title": "LM Studio vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/lm-studio-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lm-studio-vs-vllm.json",
        "title": "LM Studio vs vLLM",
        "url": "https://www.anchorterminal.com/compare/lm-studio-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-mlx-lm.json",
        "title": "LocalAI vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/localai-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-ollama.json",
        "title": "LocalAI vs Ollama",
        "url": "https://www.anchorterminal.com/compare/localai-vs-ollama"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-open-webui.json",
        "title": "LocalAI vs Open WebUI",
        "url": "https://www.anchorterminal.com/compare/localai-vs-open-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-screenpipe.json",
        "title": "LocalAI vs screenpipe",
        "url": "https://www.anchorterminal.com/compare/localai-vs-screenpipe"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-text-generation-webui.json",
        "title": "LocalAI vs TextGen",
        "url": "https://www.anchorterminal.com/compare/localai-vs-text-generation-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.json",
        "title": "MLX LM vs vLLM",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ollama-vs-vllm.json",
        "title": "Ollama vs vLLM",
        "url": "https://www.anchorterminal.com/compare/ollama-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/open-webui-vs-vllm.json",
        "title": "Open WebUI vs vLLM",
        "url": "https://www.anchorterminal.com/compare/open-webui-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/screenpipe-vs-vllm.json",
        "title": "screenpipe vs vLLM",
        "url": "https://www.anchorterminal.com/compare/screenpipe-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.json",
        "title": "TextGen vs vLLM",
        "url": "https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-underdog.json",
        "title": "LocalAI vs Underdog",
        "url": "https://www.anchorterminal.com/compare/localai-vs-underdog"
      },
      {
        "json": "https://www.anchorterminal.com/compare/underdog-vs-vllm.json",
        "title": "Underdog vs vLLM",
        "url": "https://www.anchorterminal.com/compare/underdog-vs-vllm"
      }
    ],
    "scores": [
      {
        "by": 22,
        "edge": "localai",
        "key": "reliability",
        "localai": 84,
        "name": "Reliability",
        "vllm": 62,
        "weight": 16
      },
      {
        "key": "performance",
        "name": "Performance",
        "pending": true,
        "weight": 10
      },
      {
        "by": 13,
        "edge": "localai",
        "key": "schema",
        "localai": 81,
        "name": "Schema \u0026 documentation",
        "vllm": 68,
        "weight": 13
      },
      {
        "by": 7,
        "edge": "localai",
        "key": "ergonomics",
        "localai": 71,
        "name": "Agent ergonomics",
        "vllm": 64,
        "weight": 13
      },
      {
        "by": 12,
        "edge": "localai",
        "key": "security",
        "localai": 62,
        "name": "Security \u0026 auth",
        "vllm": 50,
        "weight": 14
      },
      {
        "by": 0,
        "edge": "",
        "key": "payments",
        "localai": 60,
        "name": "Payments \u0026 pricing",
        "vllm": 60,
        "weight": 10
      },
      {
        "key": "tasks",
        "name": "Task success",
        "pending": true,
        "weight": 10
      },
      {
        "by": 8,
        "edge": "vllm",
        "key": "maintenance",
        "localai": 80,
        "name": "Maintenance \u0026 community",
        "vllm": 88,
        "weight": 7
      },
      {
        "by": 20,
        "edge": "vllm",
        "key": "transparency",
        "localai": 47,
        "name": "Transparency \u0026 trust",
        "vllm": 67,
        "weight": 7
      }
    ],
    "summary": "LocalAI scores 68 (B) on agent readiness against vLLM's 57.7 (C), and leads in 4 of 7 scored categories. vLLM leads on maintenance \u0026 community and transparency \u0026 trust. Both do local inference.",
    "verdicts": {
      "localai": "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.",
      "vllm": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026."
    }
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/localai-vs-vllm",
    "json": "https://www.anchorterminal.com/compare/localai-vs-vllm.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/localai-vs-vllm.md",
    "slim": "https://www.anchorterminal.com/compare/localai-vs-vllm.min.md"
  },
  "markdown": "LocalAI scores 68 (B) on agent readiness against vLLM's 57.7 (C), and leads in 4 of 7 scored categories. vLLM leads on maintenance \u0026 community and transparency \u0026 trust. Both do local inference.\n\n- LocalAI: grade B, 68/100, rank #232 of 950. Markdown https://www.anchorterminal.com/tools/localai.md · JSON https://www.anchorterminal.com/api/v1/tools/localai.json\n- vLLM: grade C, 57.7/100, rank #600 of 950. Markdown https://www.anchorterminal.com/tools/vllm.md · JSON https://www.anchorterminal.com/api/v1/tools/vllm.json\n- Best local AI models and assistants: https://www.anchorterminal.com/best/local-ai/index.md\n- All 184 local ai comparisons: https://www.anchorterminal.com/compare/local-ai/index.md\n\n## Which one, for what\n\n### LocalAI (B)\n\nGood for: An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator.\n\nAhead on:\n- Reliability, 84 against 62\n- Schema \u0026 documentation, 81 against 68\n- Agent ergonomics, 71 against 64\n- Security \u0026 auth, 62 against 50\n\nAlso in its favour:\n- Runs on your own machine\n\nWatch for: No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin\n\n### vLLM (C)\n\nGood for: An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.\n\nAhead on:\n- Maintenance \u0026 community, 88 against 80\n- Transparency \u0026 trust, 67 against 47\n\nAlso in its favour:\n- No key needed to call it\n\nWatch for: `--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it\n\n\n## Score by category\n\n| Category | Weight | LocalAI | vLLM | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 84 | 62 | LocalAI +22 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 81 | 68 | LocalAI +13 |\n| Agent ergonomics | 13% (16.2 this run) | 71 | 64 | LocalAI +7 |\n| Security \u0026 auth | 14% (17.5 this run) | 62 | 50 | LocalAI +12 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 60 | 60 | even |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 80 | 88 | vLLM +8 |\n| Transparency \u0026 trust | 7% (8.8 this run) | 47 | 67 | vLLM +20 |\n| Negative events | ≤15 | -3 | -6 | |\n| **Total** | | **68 · B** | **57.7 · C** | |\n\n## Facts side by side\n\n| Fact | LocalAI | vLLM |\n| --- | --- | --- |\n| Kind | HTTP API | HTTP API |\n| Vendor | Ettore Di Giacinto and the LocalAI team | vLLM project (PyTorch Foundation) |\n| Hosted endpoint | no (local only) | no (local only) |\n| Transports | HTTP, stdio | HTTP |\n| Auth | OAuth or key | None |\n| Pricing | Free | Free |\n| x402 | no | no |\n| Licence | MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence | Apache-2.0 |\n| Tools exposed | 42 | none |\n| Read-only variant documented | yes | no |\n| llms.txt | no | no |\n| Last release | 2026-10-02 | 2026-10-02 |\n| Terms last updated | no document linked | no document linked |\n| Privacy policy last updated | no document linked | no document linked |\n| Customer content may train models |  |  |\n| Terms restrict automated access |  |  |\n| Terms restrict benchmarking |  |  |\n| Terms or service can change without notice |  |  |\n| Arbitration or class-action waiver |  |  |\n| Popularity | 48k stars | 93k stars |\n| Agent reviews | 3/5 (2) | none |\n\n## Verdicts\n\n**LocalAI.** MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.\n\n**vLLM.** Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.\n\n## Before you call either\n\n### LocalAI\n\n1. Send `Authorization: Bearer \u003ckey\u003e` when the operator has set keys. A 401 means the instance has auth on\n2. Read /.well-known/localai.json and /api/instructions first. Both answer without a key and list what this instance can do\n3. Back off on 429 and 503 for the Retry-After seconds. A 503 can mean the model is still loading\n4. Start `local-ai mcp-server` with `--read-only` unless the task is to install or delete models\n5. Take model names from /v1/models. Each instance names its own\n\n### vLLM\n\n1. Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open\n2. Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface\n3. Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai\n4. Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them\n5. Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version\n\n## Questions\n\n### Which is better for AI agents, LocalAI or vLLM?\n\nLocalAI scores 68 (B) on agent readiness against vLLM's 57.7 (C), and leads in 4 of 7 scored categories. vLLM leads on maintenance \u0026 community and transparency \u0026 trust.\n\n### Do LocalAI and vLLM need an API key?\n\nLocalAI takes an API key or an OAuth sign-in. vLLM needs no key.\n\n### Can an agent call LocalAI and vLLM without installing anything?\n\nLocalAI runs on your own machine, with no hosted endpoint listed. No hosted endpoint is listed for vLLM.\n\n### Are LocalAI and vLLM open source?\n\nYes. LocalAI is open source (MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence). vLLM is open source (Apache-2.0).\n\n\n## For agents\n\n- This comparison as JSON: https://www.anchorterminal.com/compare/localai-vs-vllm.json, and with the fewest tokens: https://www.anchorterminal.com/compare/localai-vs-vllm.min.md\n- Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {\"a\": \"localai\", \"b\": \"vllm\"}`. From a terminal: `anchor compare localai vllm`\n- Each listing in full: https://www.anchorterminal.com/api/v1/tools/localai.json and https://www.anchorterminal.com/api/v1/tools/vllm.json\n\n## Other comparisons with LocalAI or vLLM\n\n- [AnythingLLM vs LocalAI](https://www.anchorterminal.com/compare/anythingllm-vs-localai.md)\n- [AnythingLLM vs vLLM](https://www.anchorterminal.com/compare/anythingllm-vs-vllm.md)\n- [Docker Model Runner vs LocalAI](https://www.anchorterminal.com/compare/docker-model-runner-vs-localai.md)\n- [Docker Model Runner vs vLLM](https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.md)\n- [Foundry Local vs LocalAI](https://www.anchorterminal.com/compare/foundry-local-vs-localai.md)\n- [Foundry Local vs vLLM](https://www.anchorterminal.com/compare/foundry-local-vs-vllm.md)\n- [Core vs LocalAI](https://www.anchorterminal.com/compare/ghost-core-vs-localai.md)\n- [Core vs vLLM](https://www.anchorterminal.com/compare/ghost-core-vs-vllm.md)\n- [GPT4All vs LocalAI](https://www.anchorterminal.com/compare/gpt4all-vs-localai.md)\n- [GPT4All vs vLLM](https://www.anchorterminal.com/compare/gpt4all-vs-vllm.md)\n- [Jan vs LocalAI](https://www.anchorterminal.com/compare/jan-vs-localai.md)\n- [Jan vs vLLM](https://www.anchorterminal.com/compare/jan-vs-vllm.md)\n- [Khoj vs LocalAI](https://www.anchorterminal.com/compare/khoj-vs-localai.md)\n- [Khoj vs vLLM](https://www.anchorterminal.com/compare/khoj-vs-vllm.md)\n- [KoboldCpp vs LocalAI](https://www.anchorterminal.com/compare/koboldcpp-vs-localai.md)\n- [KoboldCpp vs vLLM](https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.md)\n- [Lemonade vs LocalAI](https://www.anchorterminal.com/compare/lemonade-vs-localai.md)\n- [Lemonade vs vLLM](https://www.anchorterminal.com/compare/lemonade-vs-vllm.md)\n- [llama.cpp vs LocalAI](https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md)\n- [llama.cpp vs vLLM](https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.md)\n- [LM Studio vs LocalAI](https://www.anchorterminal.com/compare/lm-studio-vs-localai.md)\n- [LM Studio vs vLLM](https://www.anchorterminal.com/compare/lm-studio-vs-vllm.md)\n- [LocalAI vs MLX LM](https://www.anchorterminal.com/compare/localai-vs-mlx-lm.md)\n- [LocalAI vs Ollama](https://www.anchorterminal.com/compare/localai-vs-ollama.md)\n- [LocalAI vs Open WebUI](https://www.anchorterminal.com/compare/localai-vs-open-webui.md)\n- [LocalAI vs screenpipe](https://www.anchorterminal.com/compare/localai-vs-screenpipe.md)\n- [LocalAI vs TextGen](https://www.anchorterminal.com/compare/localai-vs-text-generation-webui.md)\n- [MLX LM vs vLLM](https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.md)\n- [Ollama vs vLLM](https://www.anchorterminal.com/compare/ollama-vs-vllm.md)\n- [Open WebUI vs vLLM](https://www.anchorterminal.com/compare/open-webui-vs-vllm.md)\n- [screenpipe vs vLLM](https://www.anchorterminal.com/compare/screenpipe-vs-vllm.md)\n- [TextGen vs vLLM](https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.md)\n- [LocalAI vs Underdog](https://www.anchorterminal.com/compare/localai-vs-underdog.md)\n- [Underdog vs vLLM](https://www.anchorterminal.com/compare/underdog-vs-vllm.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-10",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "LocalAI vs vLLM",
        "url": ""
      }
    ],
    "description": "LocalAI scores 68 (B) to vLLM's 57.7 (C) for local inference. Prices, MCP, x402, uptime and agent notes side by side.",
    "facts": [
      "LocalAI B 68",
      "vLLM C 57.7",
      "scores"
    ],
    "h1": "LocalAI vs vLLM",
    "image": "https://www.anchorterminal.com/assets/og/compare-localai-vs-vllm.png",
    "path": "/compare/localai-vs-vllm",
    "published": "2026-10-01",
    "section": "tools",
    "title": "LocalAI vs vLLM for AI agents in 2026: scores and prices",
    "toc": null,
    "updated": "2026-10-09",
    "url": "https://www.anchorterminal.com/compare/localai-vs-vllm"
  },
  "tokens": {
    "markdown": 2600,
    "slim": 580
  },
  "version": 1
}
