{
  "data": {
    "a": {
      "slug": "mlx-lm",
      "name": "MLX LM",
      "vendor": "Apple Inc.",
      "vendorUrl": "https://opensource.apple.com/projects/mlx/",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source Python package and command-line tools from Apple's MLX team for running, quantising and fine-tuning language models on Apple silicon. `mlx_lm.server` exposes a local HTTP API modelled on OpenAI's chat completions.",
      "url": "https://www.anchorterminal.com/tools/mlx-lm",
      "markdownUrl": "https://www.anchorterminal.com/tools/mlx-lm.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/mlx-lm.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/mlx-lm.json",
      "repo": "https://github.com/ml-explore/mlx-lm",
      "license": "MIT",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "pypi",
          "name": "mlx-lm"
        }
      ],
      "auth": "none",
      "authNotes": "No credential, and no option to add one. `mlx_lm.server` binds 127.0.0.1:8080 by default, and `--allowed-origins` defaults to `*`, so any origin's requests are answered. Access control is left to the network or a proxy in front (https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md).",
      "pricing": "free",
      "pricingNotes": "Free under MIT, with no account, key or card. Nothing is sold. The owner pays for the hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-08).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 7300,
        "npmWeekly": null,
        "pypiWeekly": 139915,
        "asOf": "2026-10-08"
      },
      "docsUrl": "https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md",
      "capabilities": [
        "inference.local",
        "inference.open-weights"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "python",
        "pre-1.0",
        "no-auth",
        "no-telemetry"
      ],
      "lastRelease": "2026-10-01",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 52.2,
        "grade": "D",
        "agentReady": false,
        "rank": 737,
        "ranked": true,
        "rankOf": 950,
        "categoryRank": 13,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 54,
          "maintenance": 61,
          "payments": 60,
          "reliability": 66,
          "schema": 37,
          "security": 32,
          "transparency": 66
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-08"
        },
        "negative": 0,
        "verdict": "MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production.",
        "bestFor": "An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.",
        "strengths": [
          "MIT, with no telemetry, analytics or update check found in the source",
          "Installs from PyPI (`mlx-lm` 0.32.0, Python 3.11 or later) and conda-forge, with releases published to PyPI by trusted publishing from a GitHub workflow",
          "The Build and Test workflow passed on the last eight pushes to main, with 21 test files run on a macOS runner",
          "`mlx_lm.server` binds 127.0.0.1:8080 by default, caps output at 512 tokens unless told otherwise and validates field types and ranges with a 400",
          "127 commits from 82 authors on main in the 90 days to 8 October 2026"
        ],
        "weaknesses": [
          "`mlx_lm.server` has no API key or other credential option, and `--allowed-origins` defaults to `*`",
          "A request's `model` and `adapters` fields make the server download or load any Hugging Face repository or local path, with no allow-list (open issue #1892)",
          "The docs and a start-up warning say the server is not recommended for production because it has only basic security checks",
          "No OpenAPI file or llms.txt, and `SERVER.md` leaves out `tools`, `seed`, `/health` and the error responses",
          "One PyPI release in 90 days (0.32.0 on 1 October 2026, the first since 0.31.3 on 22 April), and the version is still 0.x"
        ],
        "agentNotes": [
          "Keep `mlx_lm.server` on 127.0.0.1 and pass `--allowed-origins` with the origins you trust. There is no API key, and the default answers every origin",
          "Treat any caller as able to load any model. The `model` and `adapters` request fields accept any Hugging Face repository or local path",
          "Send `max_tokens` or `max_completion_tokens` when you need more than 512 tokens, the server default",
          "Read errors as `{\"error\": \"\u003ctext\u003e\"}` with 400 for a bad field and 404 for a model that failed to load. They are not OpenAI error objects",
          "Poll `GET /health` before the first request. It answers 503 with `unavailable` when the generation thread has stopped"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "D",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 52.2
          }
        ],
        "editorialScores": {
          "ergonomics": 54,
          "maintenance": 61,
          "payments": 60,
          "reliability": 66,
          "schema": 37,
          "security": 32,
          "transparency": 65
        },
        "provenanceScore": 67
      },
      "connect": {
        "install": "pip install mlx-lm\nmlx_lm.server --model mlx-community/Mistral-7B-Instruct-v0.3-4bit   # listens on 127.0.0.1:8080",
        "http": "curl localhost:8080/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n     \"messages\": [{\"role\": \"user\", \"content\": \"Say this is a test!\"}],\n     \"temperature\": 0.7\n   }'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/mlx-lm"
      },
      "sameCompany": [
        "apple-ads"
      ],
      "area": "models",
      "provenance": {
        "legalEntity": "Apple Inc.",
        "domain": "apple.com",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/ml-explore/mlx-lm/releases",
        "securityTxt": "valid",
        "checked": "2026-10-08",
        "notes": [
          "The `LICENSE` file reads Copyright 2023 Apple Inc., and the package author on PyPI is MLX Contributors at a group.apple.com address. The repository sits in GitHub's ml-explore organisation and has no website of its own.",
          "opensource.apple.com/projects/mlx describes the MLX framework and does not name MLX LM. Its footer links Apple's website terms and general privacy policy, which do not govern this software, so terms and privacy are left empty.",
          "www.apple.com/.well-known/security.txt is valid until 6 October 2027 and is Apple's corporate file. The repository's own policy takes reports through GitHub private vulnerability reporting.",
          "There is no shared hosted endpoint. The server runs on the owner's machine."
        ],
        "score": 67
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/mlx-lm.json",
      "live": {
        "slug": "mlx-lm",
        "versions": [
          {
            "registry": "github",
            "name": "ml-explore/mlx-lm",
            "version": "v0.31.3",
            "released": "2026-04-22",
            "seenAt": "2026-10-09T17:06:59.205890072Z"
          },
          {
            "registry": "pypi",
            "name": "mlx-lm",
            "version": "0.32.0",
            "released": "2026-10-01",
            "seenAt": "2026-10-09T17:06:59.013315523Z"
          }
        ],
        "githubStars": 7257,
        "pypiWeekly": 139648,
        "securityTxt": {
          "url": "https://apple.com/.well-known/security.txt",
          "state": "valid",
          "expires": "2027-10-06T01:00:00.000Z",
          "checkedAt": "2026-10-09T15:39:52.646042911Z"
        },
        "updatedAt": "2026-10-09T17:06:59.205890072Z"
      }
    },
    "answer": "vLLM scores 57.7 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 5 of 7 scored categories.",
    "b": {
      "slug": "vllm",
      "name": "vLLM",
      "vendor": "vLLM project (PyTorch Foundation)",
      "vendorUrl": "https://vllm.ai",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "vLLM is an open-source inference and serving engine for open-weight language models. `vllm serve` runs an HTTP server with OpenAI-compatible, Anthropic Messages, embedding, reranking and transcription routes on the owner's own GPUs or CPUs.",
      "url": "https://www.anchorterminal.com/tools/vllm",
      "markdownUrl": "https://www.anchorterminal.com/tools/vllm.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/vllm.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/vllm.json",
      "repo": "https://github.com/vllm-project/vllm",
      "license": "Apache-2.0",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "pypi",
          "name": "vllm"
        },
        {
          "registry": "oci",
          "name": "vllm/vllm-openai"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one or several keys) or `VLLM_API_KEY` turns on a Bearer check for paths under `/v1`, `/v2`, `/inference` and `/cohere` only, so `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank` and control routes such as `/pause` stay open. Keys have no scopes and change with a restart. The key is read from the `Authorization` header, never the query string. gRPC has no authentication (https://github.com/vllm-project/vllm/blob/main/docs/usage/security.md).",
      "pricing": "free",
      "pricingNotes": "Free under Apache-2.0, with no account, key or card. Nothing is sold by the project. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-09).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 93444,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-09"
      },
      "docsUrl": "https://docs.vllm.ai/en/stable/",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "speech.stt",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "telemetry-default-on"
      ],
      "lastRelease": "2026-10-02",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 57.7,
        "grade": "C",
        "agentReady": false,
        "rank": 600,
        "ranked": true,
        "rankOf": 950,
        "categoryRank": 8,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 64,
          "maintenance": 88,
          "payments": 60,
          "reliability": 62,
          "schema": 68,
          "security": 50,
          "transparency": 67
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-09"
        },
        "negative": -6,
        "negativeNotes": [
          "2026-06-02. GHSA-94f4-hr76-p5j6 (CVE-2026-48746, 9.1), a crafted Host header bypassed the API key check on the OpenAI routes, fixed in 0.22.0. With GHSA-4r2x-xpjr-7cvv (CVE-2026-22778, 9.8) of 2 February 2026, code execution through video decoding fixed in 0.14.1, these are the two critical advisories of the last 12 months. Both were fixed and published with CVEs, so they decay, -3. https://github.com/vllm-project/vllm/security/advisories/GHSA-94f4-hr76-p5j6; https://github.com/vllm-project/vllm/security/advisories/GHSA-4r2x-xpjr-7cvv",
          "2026-10-06. GHSA-h3rc-6mm3-gc2m (8.1), a request field could select the processor code a server started with `--trust-remote-code` imports, fixed in 0.31.0, one of 50 advisories published since 11 July 2026 (10 high, 36 medium, 4 low), most of them requests that crash or exhaust the engine. All name a fixed version, and eleven were published on 9 October 2026 months after their fixes, -3. https://github.com/vllm-project/vllm/security/advisories/GHSA-h3rc-6mm3-gc2m; https://github.com/vllm-project/vllm/security/advisories"
        ],
        "verdict": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.",
        "bestFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "strengths": [
          "OpenAI chat, completions, responses and embeddings, Anthropic `/v1/messages`, Cohere embed and rerank, transcription and `/v1/systemone` from one server",
          "Apache-2.0, with a written three-stage deprecation policy and release notes that carry a breaking changes section",
          "Eight stable releases between 12 July and 2 October 2026, and v0.31.0 lists 717 commits from 307 contributors",
          "A 650-line security guide names every route the API key does and does not protect, and the limits of multi-tenant use",
          "Usage statistics are documented field by field, with `VLLM_NO_USAGE_STATS`, `DO_NOT_TRACK` or a file as opt-outs"
        ],
        "weaknesses": [
          "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it",
          "No key by default, the server binds every interface when `--host` is unset, and CORS allows any origin",
          "At least 81 GitHub security advisories in 12 months, two critical, most of them remote crashes or resource exhaustion",
          "Pre-1.0 (0.31.0), with breaking changes in each fortnightly release and compatibility kept for a limited number of minor versions",
          "Usage statistics are sent to stats.vllm.ai by default, and no privacy policy or retention period for them was found"
        ],
        "agentNotes": [
          "Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open",
          "Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface",
          "Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai",
          "Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them",
          "Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 57.7
          }
        ],
        "editorialScores": {
          "ergonomics": 64,
          "maintenance": 88,
          "payments": 60,
          "reliability": 62,
          "schema": 68,
          "security": 50,
          "transparency": 80
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "uv pip install vllm --torch-backend=auto\nvllm serve Qwen/Qwen2.5-1.5B-Instruct   # listens on port 8000",
        "http": "curl http://localhost:8000/v1/chat/completions \\\n    -H \"Content-Type: application/json\" \\\n    -d '{\n        \"model\": \"Qwen/Qwen2.5-1.5B-Instruct\",\n        \"messages\": [\n            {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n            {\"role\": \"user\", \"content\": \"Who won the world series in 2020?\"}\n        ]\n    }'",
        "claudeCode": "ANTHROPIC_BASE_URL=http://localhost:8000 \\\nANTHROPIC_API_KEY=dummy \\\nANTHROPIC_AUTH_TOKEN=dummy \\\nANTHROPIC_DEFAULT_OPUS_MODEL=my-model \\\nANTHROPIC_DEFAULT_SONNET_MODEL=my-model \\\nANTHROPIC_DEFAULT_HAIKU_MODEL=my-model \\\nclaude"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/vllm"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "The Linux Foundation (vLLM is a PyTorch Foundation project)",
        "domain": "vllm.ai",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/vllm-project/vllm/releases",
        "securityTxt": "none",
        "checked": "2026-10-09",
        "notes": [
          "vllm.ai links no terms and no privacy policy, and its footer reads © 2026 vLLM. The project publishes none for the software or for stats.vllm.ai, so the Apache-2.0 licence stands in for terms.",
          "pytorch.org/projects/vllm/ lists vLLM among PyTorch Foundation projects and says UC Berkeley contributed it to the Linux Foundation in July 2024. The Linux Foundation's policies are linked from that page and are not specific to vLLM.",
          "vllm.ai/.well-known/security.txt and docs.vllm.ai/.well-known/security.txt return 404. SECURITY.md asks for private reports through GitHub.",
          "There's no shared hosted endpoint. The server runs on the owner's hardware. The software posts usage statistics to stats.vllm.ai unless turned off."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/vllm.json",
      "live": {
        "slug": "vllm",
        "versions": [
          {
            "registry": "github",
            "name": "vllm-project/vllm",
            "version": "v0.31.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:27:30.444059802Z"
          },
          {
            "registry": "pypi",
            "name": "vllm",
            "version": "0.31.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-09T17:27:30.259454009Z"
          }
        ],
        "githubStars": 93457,
        "pypiWeekly": 444466,
        "updatedAt": "2026-10-09T17:27:30.444059802Z"
      }
    },
    "facts": [
      {
        "a": "HTTP API",
        "b": "HTTP API",
        "name": "Kind"
      },
      {
        "a": "Apple Inc.",
        "b": "vLLM project (PyTorch Foundation)",
        "name": "Vendor"
      },
      {
        "a": "no (local only)",
        "b": "no (local only)",
        "name": "Hosted endpoint"
      },
      {
        "a": "HTTP",
        "b": "HTTP",
        "name": "Transports"
      },
      {
        "a": "None",
        "b": "None",
        "name": "Auth"
      },
      {
        "a": "Free",
        "b": "Free",
        "name": "Pricing"
      },
      {
        "a": "no",
        "b": "no",
        "name": "x402"
      },
      {
        "a": "MIT",
        "b": "Apache-2.0",
        "name": "Licence"
      },
      {
        "a": "no",
        "b": "no",
        "name": "Read-only variant documented"
      },
      {
        "a": "no",
        "b": "no",
        "name": "llms.txt"
      },
      {
        "a": "2026-10-01",
        "b": "2026-10-02",
        "name": "Last release"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Terms last updated"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Privacy policy last updated"
      },
      {
        "a": "",
        "b": "",
        "name": "Customer content may train models"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict automated access"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict benchmarking"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms or service can change without notice"
      },
      {
        "a": "",
        "b": "",
        "name": "Arbitration or class-action waiver"
      },
      {
        "a": "7.3k stars, 140k PyPI/wk",
        "b": "93k stars",
        "name": "Popularity"
      }
    ],
    "faq": [
      {
        "answer": "vLLM scores 57.7 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 5 of 7 scored categories.",
        "question": "Which is better for AI agents, MLX LM or vLLM?"
      },
      {
        "answer": "Neither needs a key.",
        "question": "Do MLX LM and vLLM need an API key?"
      },
      {
        "answer": "No hosted endpoint is listed for MLX LM. No hosted endpoint is listed for vLLM.",
        "question": "Can an agent call MLX LM and vLLM without installing anything?"
      },
      {
        "answer": "Yes. MLX LM is open source (MIT). vLLM is open source (Apache-2.0).",
        "question": "Are MLX LM and vLLM open source?"
      }
    ],
    "goodFor": [
      {
        "aheadOn": null,
        "also": [
          "No incidents deducted, where vLLM loses 6 points for them"
        ],
        "goodFor": "An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.",
        "slug": "mlx-lm",
        "watchFor": "`mlx_lm.server` has no API key or other credential option, and `--allowed-origins` defaults to `*`"
      },
      {
        "aheadOn": [
          "Schema \u0026 documentation, 68 against 37",
          "Agent ergonomics, 64 against 54",
          "Security \u0026 auth, 50 against 32",
          "Maintenance \u0026 community, 88 against 61"
        ],
        "also": null,
        "goodFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "slug": "vllm",
        "watchFor": "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it"
      }
    ],
    "job": {
      "capability": "inference.local",
      "name": "Local inference"
    },
    "others": [
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-mlx-lm.json",
        "title": "AnythingLLM vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-vllm.json",
        "title": "AnythingLLM vs vLLM",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-mlx-lm.json",
        "title": "Docker Model Runner vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.json",
        "title": "Docker Model Runner vs vLLM",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm.json",
        "title": "Foundry Local vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-vllm.json",
        "title": "Foundry Local vs vLLM",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-mlx-lm.json",
        "title": "Core vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-vllm.json",
        "title": "Core vs vLLM",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-mlx-lm.json",
        "title": "GPT4All vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-vllm.json",
        "title": "GPT4All vs vLLM",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-mlx-lm.json",
        "title": "Jan vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/jan-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-vllm.json",
        "title": "Jan vs vLLM",
        "url": "https://www.anchorterminal.com/compare/jan-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-mlx-lm.json",
        "title": "Khoj vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-vllm.json",
        "title": "Khoj vs vLLM",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-mlx-lm.json",
        "title": "KoboldCpp vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.json",
        "title": "KoboldCpp vs vLLM",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-mlx-lm.json",
        "title": "Lemonade vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-vllm.json",
        "title": "Lemonade vs vLLM",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.json",
        "title": "llama.cpp vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.json",
        "title": "llama.cpp vs vLLM",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lm-studio-vs-mlx-lm.json",
        "title": "LM Studio vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/lm-studio-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lm-studio-vs-vllm.json",
        "title": "LM Studio vs vLLM",
        "url": "https://www.anchorterminal.com/compare/lm-studio-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-mlx-lm.json",
        "title": "LocalAI vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/localai-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-vllm.json",
        "title": "LocalAI vs vLLM",
        "url": "https://www.anchorterminal.com/compare/localai-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-ollama.json",
        "title": "MLX LM vs Ollama",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-ollama"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-open-webui.json",
        "title": "MLX LM vs Open WebUI",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-open-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-screenpipe.json",
        "title": "MLX LM vs screenpipe",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-screenpipe"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-text-generation-webui.json",
        "title": "MLX LM vs TextGen",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-text-generation-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ollama-vs-vllm.json",
        "title": "Ollama vs vLLM",
        "url": "https://www.anchorterminal.com/compare/ollama-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/open-webui-vs-vllm.json",
        "title": "Open WebUI vs vLLM",
        "url": "https://www.anchorterminal.com/compare/open-webui-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/screenpipe-vs-vllm.json",
        "title": "screenpipe vs vLLM",
        "url": "https://www.anchorterminal.com/compare/screenpipe-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.json",
        "title": "TextGen vs vLLM",
        "url": "https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-underdog.json",
        "title": "MLX LM vs Underdog",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-underdog"
      },
      {
        "json": "https://www.anchorterminal.com/compare/underdog-vs-vllm.json",
        "title": "Underdog vs vLLM",
        "url": "https://www.anchorterminal.com/compare/underdog-vs-vllm"
      }
    ],
    "scores": [
      {
        "by": 4,
        "edge": "mlx-lm",
        "key": "reliability",
        "mlx-lm": 66,
        "name": "Reliability",
        "vllm": 62,
        "weight": 16
      },
      {
        "key": "performance",
        "name": "Performance",
        "pending": true,
        "weight": 10
      },
      {
        "by": 31,
        "edge": "vllm",
        "key": "schema",
        "mlx-lm": 37,
        "name": "Schema \u0026 documentation",
        "vllm": 68,
        "weight": 13
      },
      {
        "by": 10,
        "edge": "vllm",
        "key": "ergonomics",
        "mlx-lm": 54,
        "name": "Agent ergonomics",
        "vllm": 64,
        "weight": 13
      },
      {
        "by": 18,
        "edge": "vllm",
        "key": "security",
        "mlx-lm": 32,
        "name": "Security \u0026 auth",
        "vllm": 50,
        "weight": 14
      },
      {
        "by": 0,
        "edge": "",
        "key": "payments",
        "mlx-lm": 60,
        "name": "Payments \u0026 pricing",
        "vllm": 60,
        "weight": 10
      },
      {
        "key": "tasks",
        "name": "Task success",
        "pending": true,
        "weight": 10
      },
      {
        "by": 27,
        "edge": "vllm",
        "key": "maintenance",
        "mlx-lm": 61,
        "name": "Maintenance \u0026 community",
        "vllm": 88,
        "weight": 7
      },
      {
        "by": 1,
        "edge": "vllm",
        "key": "transparency",
        "mlx-lm": 66,
        "name": "Transparency \u0026 trust",
        "vllm": 67,
        "weight": 7
      }
    ],
    "summary": "vLLM scores 57.7 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 5 of 7 scored categories. Both do local inference.",
    "verdicts": {
      "mlx-lm": "MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production.",
      "vllm": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026."
    }
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm",
    "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.md",
    "slim": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.min.md"
  },
  "markdown": "vLLM scores 57.7 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 5 of 7 scored categories. Both do local inference.\n\n- MLX LM: grade D, 52.2/100, rank #737 of 950. Markdown https://www.anchorterminal.com/tools/mlx-lm.md · JSON https://www.anchorterminal.com/api/v1/tools/mlx-lm.json\n- vLLM: grade C, 57.7/100, rank #600 of 950. Markdown https://www.anchorterminal.com/tools/vllm.md · JSON https://www.anchorterminal.com/api/v1/tools/vllm.json\n- Best local AI models and assistants: https://www.anchorterminal.com/best/local-ai/index.md\n- All 184 local ai comparisons: https://www.anchorterminal.com/compare/local-ai/index.md\n\n## Which one, for what\n\n### MLX LM (D)\n\nGood for: An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.\n\nAlso in its favour:\n- No incidents deducted, where vLLM loses 6 points for them\n\nWatch for: `mlx_lm.server` has no API key or other credential option, and `--allowed-origins` defaults to `*`\n\n### vLLM (C)\n\nGood for: An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.\n\nAhead on:\n- Schema \u0026 documentation, 68 against 37\n- Agent ergonomics, 64 against 54\n- Security \u0026 auth, 50 against 32\n- Maintenance \u0026 community, 88 against 61\n\nWatch for: `--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it\n\n\n## Score by category\n\n| Category | Weight | MLX LM | vLLM | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 66 | 62 | MLX LM +4 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 37 | 68 | vLLM +31 |\n| Agent ergonomics | 13% (16.2 this run) | 54 | 64 | vLLM +10 |\n| Security \u0026 auth | 14% (17.5 this run) | 32 | 50 | vLLM +18 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 60 | 60 | even |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 61 | 88 | vLLM +27 |\n| Transparency \u0026 trust | 7% (8.8 this run) | 66 | 67 | vLLM +1 |\n| Negative events | ≤15 | 0 | -6 | |\n| **Total** | | **52.2 · D** | **57.7 · C** | |\n\n## Facts side by side\n\n| Fact | MLX LM | vLLM |\n| --- | --- | --- |\n| Kind | HTTP API | HTTP API |\n| Vendor | Apple Inc. | vLLM project (PyTorch Foundation) |\n| Hosted endpoint | no (local only) | no (local only) |\n| Transports | HTTP | HTTP |\n| Auth | None | None |\n| Pricing | Free | Free |\n| x402 | no | no |\n| Licence | MIT | Apache-2.0 |\n| Read-only variant documented | no | no |\n| llms.txt | no | no |\n| Last release | 2026-10-01 | 2026-10-02 |\n| Terms last updated | no document linked | no document linked |\n| Privacy policy last updated | no document linked | no document linked |\n| Customer content may train models |  |  |\n| Terms restrict automated access |  |  |\n| Terms restrict benchmarking |  |  |\n| Terms or service can change without notice |  |  |\n| Arbitration or class-action waiver |  |  |\n| Popularity | 7.3k stars, 140k PyPI/wk | 93k stars |\n\n## Verdicts\n\n**MLX LM.** MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production.\n\n**vLLM.** Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.\n\n## Before you call either\n\n### MLX LM\n\n1. Keep `mlx_lm.server` on 127.0.0.1 and pass `--allowed-origins` with the origins you trust. There is no API key, and the default answers every origin\n2. Treat any caller as able to load any model. The `model` and `adapters` request fields accept any Hugging Face repository or local path\n3. Send `max_tokens` or `max_completion_tokens` when you need more than 512 tokens, the server default\n4. Read errors as `{\"error\": \"\u003ctext\u003e\"}` with 400 for a bad field and 404 for a model that failed to load. They are not OpenAI error objects\n5. Poll `GET /health` before the first request. It answers 503 with `unavailable` when the generation thread has stopped\n\n### vLLM\n\n1. Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open\n2. Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface\n3. Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai\n4. Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them\n5. Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version\n\n## Questions\n\n### Which is better for AI agents, MLX LM or vLLM?\n\nvLLM scores 57.7 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 5 of 7 scored categories.\n\n### Do MLX LM and vLLM need an API key?\n\nNeither needs a key.\n\n### Can an agent call MLX LM and vLLM without installing anything?\n\nNo hosted endpoint is listed for MLX LM. No hosted endpoint is listed for vLLM.\n\n### Are MLX LM and vLLM open source?\n\nYes. MLX LM is open source (MIT). vLLM is open source (Apache-2.0).\n\n\n## For agents\n\n- This comparison as JSON: https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.json, and with the fewest tokens: https://www.anchorterminal.com/compare/mlx-lm-vs-vllm.min.md\n- Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {\"a\": \"mlx-lm\", \"b\": \"vllm\"}`. From a terminal: `anchor compare mlx-lm vllm`\n- Each listing in full: https://www.anchorterminal.com/api/v1/tools/mlx-lm.json and https://www.anchorterminal.com/api/v1/tools/vllm.json\n\n## Other comparisons with MLX LM or vLLM\n\n- [AnythingLLM vs MLX LM](https://www.anchorterminal.com/compare/anythingllm-vs-mlx-lm.md)\n- [AnythingLLM vs vLLM](https://www.anchorterminal.com/compare/anythingllm-vs-vllm.md)\n- [Docker Model Runner vs MLX LM](https://www.anchorterminal.com/compare/docker-model-runner-vs-mlx-lm.md)\n- [Docker Model Runner vs vLLM](https://www.anchorterminal.com/compare/docker-model-runner-vs-vllm.md)\n- [Foundry Local vs MLX LM](https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm.md)\n- [Foundry Local vs vLLM](https://www.anchorterminal.com/compare/foundry-local-vs-vllm.md)\n- [Core vs MLX LM](https://www.anchorterminal.com/compare/ghost-core-vs-mlx-lm.md)\n- [Core vs vLLM](https://www.anchorterminal.com/compare/ghost-core-vs-vllm.md)\n- [GPT4All vs MLX LM](https://www.anchorterminal.com/compare/gpt4all-vs-mlx-lm.md)\n- [GPT4All vs vLLM](https://www.anchorterminal.com/compare/gpt4all-vs-vllm.md)\n- [Jan vs MLX LM](https://www.anchorterminal.com/compare/jan-vs-mlx-lm.md)\n- [Jan vs vLLM](https://www.anchorterminal.com/compare/jan-vs-vllm.md)\n- [Khoj vs MLX LM](https://www.anchorterminal.com/compare/khoj-vs-mlx-lm.md)\n- [Khoj vs vLLM](https://www.anchorterminal.com/compare/khoj-vs-vllm.md)\n- [KoboldCpp vs MLX LM](https://www.anchorterminal.com/compare/koboldcpp-vs-mlx-lm.md)\n- [KoboldCpp vs vLLM](https://www.anchorterminal.com/compare/koboldcpp-vs-vllm.md)\n- [Lemonade vs MLX LM](https://www.anchorterminal.com/compare/lemonade-vs-mlx-lm.md)\n- [Lemonade vs vLLM](https://www.anchorterminal.com/compare/lemonade-vs-vllm.md)\n- [llama.cpp vs MLX LM](https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.md)\n- [llama.cpp vs vLLM](https://www.anchorterminal.com/compare/llama-cpp-vs-vllm.md)\n- [LM Studio vs MLX LM](https://www.anchorterminal.com/compare/lm-studio-vs-mlx-lm.md)\n- [LM Studio vs vLLM](https://www.anchorterminal.com/compare/lm-studio-vs-vllm.md)\n- [LocalAI vs MLX LM](https://www.anchorterminal.com/compare/localai-vs-mlx-lm.md)\n- [LocalAI vs vLLM](https://www.anchorterminal.com/compare/localai-vs-vllm.md)\n- [MLX LM vs Ollama](https://www.anchorterminal.com/compare/mlx-lm-vs-ollama.md)\n- [MLX LM vs Open WebUI](https://www.anchorterminal.com/compare/mlx-lm-vs-open-webui.md)\n- [MLX LM vs screenpipe](https://www.anchorterminal.com/compare/mlx-lm-vs-screenpipe.md)\n- [MLX LM vs TextGen](https://www.anchorterminal.com/compare/mlx-lm-vs-text-generation-webui.md)\n- [Ollama vs vLLM](https://www.anchorterminal.com/compare/ollama-vs-vllm.md)\n- [Open WebUI vs vLLM](https://www.anchorterminal.com/compare/open-webui-vs-vllm.md)\n- [screenpipe vs vLLM](https://www.anchorterminal.com/compare/screenpipe-vs-vllm.md)\n- [TextGen vs vLLM](https://www.anchorterminal.com/compare/text-generation-webui-vs-vllm.md)\n- [MLX LM vs Underdog](https://www.anchorterminal.com/compare/mlx-lm-vs-underdog.md)\n- [Underdog vs vLLM](https://www.anchorterminal.com/compare/underdog-vs-vllm.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-10",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "MLX LM vs vLLM",
        "url": ""
      }
    ],
    "description": "vLLM scores 57.7 (C) to MLX LM's 52.2 (D) for local inference. Prices, MCP, x402, uptime and agent notes side by side.",
    "facts": [
      "MLX LM D 52.2",
      "vLLM C 57.7",
      "scores"
    ],
    "h1": "MLX LM vs vLLM",
    "image": "https://www.anchorterminal.com/assets/og/compare-mlx-lm-vs-vllm.png",
    "path": "/compare/mlx-lm-vs-vllm",
    "published": "2026-10-01",
    "section": "tools",
    "title": "MLX LM vs vLLM for AI agents in 2026: scores and prices",
    "toc": null,
    "updated": "2026-10-09",
    "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-vllm"
  },
  "tokens": {
    "markdown": 2450,
    "slim": 480
  },
  "version": 1
}
