{
  "data": {
    "a": {
      "slug": "llama-cpp",
      "name": "llama.cpp",
      "vendor": "ggml.ai (Hugging Face)",
      "vendorUrl": "https://llama.app",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source C/C++ engine for running GGUF models locally, with a web interface and compatible model APIs.",
      "url": "https://www.anchorterminal.com/tools/llama-cpp",
      "markdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/llama-cpp.json",
      "repo": "https://github.com/ggml-org/llama.cpp",
      "license": "MIT",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "ghcr.io/ggml-org/llama.cpp"
        },
        {
          "registry": "pypi",
          "name": "gguf"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one key or a comma-separated list) or `--api-key-file` (one key a line) turns on a check for every route but /health and the web UI's files, with the key sent as `Authorization: Bearer` or `X-Api-Key`, never in the query string. Keys have no scopes and change only with a restart. TLS is built in with `--ssl-key-file` and `--ssl-cert-file`. The server binds 127.0.0.1:8080 by default, and CORS reflects any Origin with credentials allowed unless built-in tools, MCP servers or `--agent` are on, when it narrows to localhost (https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).",
      "pricing": "free",
      "pricingNotes": "Free under MIT, with no account, key or card. Nothing is sold. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 130200,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "no-telemetry"
      ],
      "lastRelease": "2026-09-23",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 60.2,
        "grade": "C",
        "agentReady": false,
        "rank": 476,
        "ranked": true,
        "rankOf": 842,
        "categoryRank": 6,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 60
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -1,
        "negativeNotes": [
          "2026-03-26. GHSA-j8rj-fmpv-wcxw (CVE-2026-34159, 9.8 at NVD), unauthenticated code execution through a GRAPH_COMPUTE bypass in the RPC backend, the most serious of four advisories published between January and March 2026 (the others a llama-server out-of-bounds write through a negative `n_discard` and two GGUF integer overflows). All were fixed in named builds and published as advisories, SECURITY.md says not to expose the RPC server or llama-server to untrusted networks, and the newest is more than six months old, -1. https://github.com/ggml-org/llama.cpp/security/advisories/GHSA-j8rj-fmpv-wcxw; https://github.com/ggml-org/llama.cpp/security"
        ],
        "verdict": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
        "bestFor": "An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.",
        "strengths": [
          "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads",
          "OpenAI chat completions, responses and embeddings, Anthropic messages, reranking and /v1/systemone from one server",
          "`response_fields`, `json_schema` and `grammar` control the size and shape of output, and errors carry an OpenAI-style type and code",
          "1,005 nightly builds and eight semver releases in 90 days, with 37 workflows running on every push to master",
          "Ten published GitHub advisories with CVEs and fixed builds, and SECURITY.md guidance on untrusted models and inputs"
        ],
        "weaknesses": [
          "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost",
          "No OpenAPI file of its own, and the REST API changelog stops at b4599",
          "Private security disclosure disabled since 1 June 2026, with fixes asked for as public pull requests",
          "Pre-1.0 (0.5.0), and semver releases are bare tags with no notes",
          "No official client library, and `n_predict` defaults to unlimited"
        ],
        "agentNotes": [
          "Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default",
          "Pass `n_predict` or `max_tokens`. Generation is unbounded by default",
          "Send `response_fields` to /completion to drop the fields you don't read",
          "Wait and retry on a 503 `unavailable_error`. The model is still loading",
          "Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 2.5,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 60.2
          }
        ],
        "editorialScores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 66
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "curl -LsSf https://llama.app/install.sh | sh   # or: brew install llama.cpp; winget install llama.cpp\nllama serve -hf ggml-org/Qwen3.5-0.8B-GGUF   # listens on 127.0.0.1:8080",
        "http": "curl --request POST \\\n    --url http://localhost:8080/completion \\\n    --header \"Content-Type: application/json\" \\\n    --data '{\"prompt\": \"Building a website can be done in 10 simple steps:\",\"n_predict\": 128}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/llama-cpp"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "ggml.ai, part of Hugging Face since 2026",
        "domain": "llama.app",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/ggml-org/llama.cpp/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "The repository's About link is llama.app, which says it's by the llama.cpp team and Hugging Face and links no terms, privacy or security page. ggml.ai says the company was acquired by Hugging Face in 2026 and names no address.",
          "The `LICENSE` file reads Copyright (c) 2023-2026 The ggml authors.",
          "llama.app/.well-known/security.txt and llama.app/llms.txt return 404. SECURITY.md points to GitHub private advisories while saying private disclosure is disabled.",
          "There's no shared hosted endpoint. The server runs on the owner's machine."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/llama-cpp.json",
      "live": {
        "slug": "llama-cpp",
        "versions": [
          {
            "registry": "github",
            "name": "ggml-org/llama.cpp",
            "version": "v0.6.0",
            "released": "2026-10-05",
            "seenAt": "2026-10-08T16:19:08.340661659Z"
          },
          {
            "registry": "pypi",
            "name": "gguf",
            "version": "0.19.0",
            "released": "2026-05-06",
            "seenAt": "2026-10-08T16:19:08.217189108Z"
          }
        ],
        "githubStars": 130684,
        "pypiWeekly": 1220883,
        "securityTxt": {
          "url": "https://llama.app/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-08T15:38:54.37968551Z"
        },
        "domain": {
          "domain": "llama.app",
          "registered": "2018-07-18",
          "source": "https://pubapi.registry.google/rdap/domain/llama.app",
          "checkedAt": "2026-10-04T13:04:03.05886804Z"
        },
        "updatedAt": "2026-10-08T16:19:08.340661659Z"
      }
    },
    "answer": "llama.cpp scores 60.2 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 4 of 7 scored categories. MLX LM leads on transparency \u0026 trust.",
    "b": {
      "slug": "mlx-lm",
      "name": "MLX LM",
      "vendor": "Apple Inc.",
      "vendorUrl": "https://opensource.apple.com/projects/mlx/",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source Python package and command-line tools from Apple's MLX team for running, quantising and fine-tuning language models on Apple silicon. `mlx_lm.server` exposes a local HTTP API modelled on OpenAI's chat completions.",
      "url": "https://www.anchorterminal.com/tools/mlx-lm",
      "markdownUrl": "https://www.anchorterminal.com/tools/mlx-lm.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/mlx-lm.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/mlx-lm.json",
      "repo": "https://github.com/ml-explore/mlx-lm",
      "license": "MIT",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "pypi",
          "name": "mlx-lm"
        }
      ],
      "auth": "none",
      "authNotes": "No credential, and no option to add one. `mlx_lm.server` binds 127.0.0.1:8080 by default, and `--allowed-origins` defaults to `*`, so any origin's requests are answered. Access control is left to the network or a proxy in front (https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md).",
      "pricing": "free",
      "pricingNotes": "Free under MIT, with no account, key or card. Nothing is sold. The owner pays for the hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-08).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 7300,
        "npmWeekly": null,
        "pypiWeekly": 139915,
        "asOf": "2026-10-08"
      },
      "docsUrl": "https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md",
      "capabilities": [
        "inference.local",
        "inference.open-weights"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "python",
        "pre-1.0",
        "no-auth",
        "no-telemetry"
      ],
      "lastRelease": "2026-10-01",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 52.2,
        "grade": "D",
        "agentReady": false,
        "rank": 657,
        "ranked": true,
        "rankOf": 842,
        "categoryRank": 12,
        "methodology": "0.4",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 54,
          "maintenance": 61,
          "payments": 60,
          "reliability": 66,
          "schema": 37,
          "security": 32,
          "transparency": 66
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-08"
        },
        "negative": 0,
        "verdict": "MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production.",
        "bestFor": "An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.",
        "strengths": [
          "MIT, with no telemetry, analytics or update check found in the source",
          "Installs from PyPI (`mlx-lm` 0.32.0, Python 3.11 or later) and conda-forge, with releases published to PyPI by trusted publishing from a GitHub workflow",
          "The Build and Test workflow passed on the last eight pushes to main, with 21 test files run on a macOS runner",
          "`mlx_lm.server` binds 127.0.0.1:8080 by default, caps output at 512 tokens unless told otherwise and validates field types and ranges with a 400",
          "127 commits from 82 authors on main in the 90 days to 8 October 2026"
        ],
        "weaknesses": [
          "`mlx_lm.server` has no API key or other credential option, and `--allowed-origins` defaults to `*`",
          "A request's `model` and `adapters` fields make the server download or load any Hugging Face repository or local path, with no allow-list (open issue #1892)",
          "The docs and a start-up warning say the server is not recommended for production because it has only basic security checks",
          "No OpenAPI file or llms.txt, and `SERVER.md` leaves out `tools`, `seed`, `/health` and the error responses",
          "One PyPI release in 90 days (0.32.0 on 1 October 2026, the first since 0.31.3 on 22 April), and the version is still 0.x"
        ],
        "agentNotes": [
          "Keep `mlx_lm.server` on 127.0.0.1 and pass `--allowed-origins` with the origins you trust. There is no API key, and the default answers every origin",
          "Treat any caller as able to load any model. The `model` and `adapters` request fields accept any Hugging Face repository or local path",
          "Send `max_tokens` or `max_completion_tokens` when you need more than 512 tokens, the server default",
          "Read errors as `{\"error\": \"\u003ctext\u003e\"}` with 400 for a bad field and 404 for a model that failed to load. They are not OpenAI error objects",
          "Poll `GET /health` before the first request. It answers 503 with `unavailable` when the generation thread has stopped"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 0,
        "avgRating": 0,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "D",
            "methodology": "0.4",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 52.2
          }
        ],
        "editorialScores": {
          "ergonomics": 54,
          "maintenance": 61,
          "payments": 60,
          "reliability": 66,
          "schema": 37,
          "security": 32,
          "transparency": 65
        },
        "provenanceScore": 67
      },
      "connect": {
        "install": "pip install mlx-lm\nmlx_lm.server --model mlx-community/Mistral-7B-Instruct-v0.3-4bit   # listens on 127.0.0.1:8080",
        "http": "curl localhost:8080/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n     \"messages\": [{\"role\": \"user\", \"content\": \"Say this is a test!\"}],\n     \"temperature\": 0.7\n   }'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/mlx-lm"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "Apple Inc.",
        "domain": "apple.com",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/ml-explore/mlx-lm/releases",
        "securityTxt": "valid",
        "checked": "2026-10-08",
        "notes": [
          "The `LICENSE` file reads Copyright 2023 Apple Inc., and the package author on PyPI is MLX Contributors at a group.apple.com address. The repository sits in GitHub's ml-explore organisation and has no website of its own.",
          "opensource.apple.com/projects/mlx describes the MLX framework and does not name MLX LM. Its footer links Apple's website terms and general privacy policy, which do not govern this software, so terms and privacy are left empty.",
          "www.apple.com/.well-known/security.txt is valid until 6 October 2027 and is Apple's corporate file. The repository's own policy takes reports through GitHub private vulnerability reporting.",
          "There is no shared hosted endpoint. The server runs on the owner's machine."
        ],
        "score": 67
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/mlx-lm.json"
    },
    "facts": [
      {
        "a": "HTTP API",
        "b": "HTTP API",
        "name": "Kind"
      },
      {
        "a": "ggml.ai (Hugging Face)",
        "b": "Apple Inc.",
        "name": "Vendor"
      },
      {
        "a": "no (local only)",
        "b": "no (local only)",
        "name": "Hosted endpoint"
      },
      {
        "a": "HTTP",
        "b": "HTTP",
        "name": "Transports"
      },
      {
        "a": "None",
        "b": "None",
        "name": "Auth"
      },
      {
        "a": "Free",
        "b": "Free",
        "name": "Pricing"
      },
      {
        "a": "no",
        "b": "no",
        "name": "x402"
      },
      {
        "a": "MIT",
        "b": "MIT",
        "name": "Licence"
      },
      {
        "a": "no",
        "b": "no",
        "name": "Read-only variant documented"
      },
      {
        "a": "no",
        "b": "no",
        "name": "llms.txt"
      },
      {
        "a": "2026-09-23",
        "b": "2026-10-01",
        "name": "Last release"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Terms last updated"
      },
      {
        "a": "no document linked",
        "b": "no document linked",
        "name": "Privacy policy last updated"
      },
      {
        "a": "",
        "b": "",
        "name": "Customer content may train models"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict automated access"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms restrict benchmarking"
      },
      {
        "a": "",
        "b": "",
        "name": "Terms or service can change without notice"
      },
      {
        "a": "",
        "b": "",
        "name": "Arbitration or class-action waiver"
      },
      {
        "a": "130k stars",
        "b": "7.3k stars, 140k PyPI/wk",
        "name": "Popularity"
      },
      {
        "a": "2.5/5 (2)",
        "b": "none",
        "name": "Agent reviews"
      }
    ],
    "faq": [
      {
        "answer": "llama.cpp scores 60.2 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 4 of 7 scored categories. MLX LM leads on transparency \u0026 trust.",
        "question": "Which is better for AI agents, llama.cpp or MLX LM?"
      },
      {
        "answer": "Neither needs a key.",
        "question": "Do llama.cpp and MLX LM need an API key?"
      },
      {
        "answer": "No hosted endpoint is listed for llama.cpp. No hosted endpoint is listed for MLX LM.",
        "question": "Can an agent call llama.cpp and MLX LM without installing anything?"
      },
      {
        "answer": "Yes. llama.cpp is open source (MIT). MLX LM is open source (MIT).",
        "question": "Are llama.cpp and MLX LM open source?"
      }
    ],
    "goodFor": [
      {
        "aheadOn": [
          "Schema \u0026 documentation, 47 against 37",
          "Agent ergonomics, 73 against 54",
          "Security \u0026 auth, 52 against 32",
          "Maintenance \u0026 community, 81 against 61"
        ],
        "also": null,
        "goodFor": "An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.",
        "slug": "llama-cpp",
        "watchFor": "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost"
      },
      {
        "aheadOn": [
          "Transparency \u0026 trust, 66 against 60"
        ],
        "also": null,
        "goodFor": "An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.",
        "slug": "mlx-lm",
        "watchFor": "`mlx_lm.server` has no API key or other credential option, and `--allowed-origins` defaults to `*`"
      }
    ],
    "job": {
      "capability": "inference.local",
      "name": "Local inference"
    },
    "others": [
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.json",
        "title": "AnythingLLM vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/anythingllm-vs-mlx-lm.json",
        "title": "AnythingLLM vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/anythingllm-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp.json",
        "title": "Docker Model Runner vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/docker-model-runner-vs-mlx-lm.json",
        "title": "Docker Model Runner vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/docker-model-runner-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp.json",
        "title": "Foundry Local vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm.json",
        "title": "Foundry Local vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp.json",
        "title": "Core vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/ghost-core-vs-mlx-lm.json",
        "title": "Core vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/ghost-core-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.json",
        "title": "GPT4All vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/gpt4all-vs-mlx-lm.json",
        "title": "GPT4All vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/gpt4all-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-llama-cpp.json",
        "title": "Jan vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/jan-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/jan-vs-mlx-lm.json",
        "title": "Jan vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/jan-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.json",
        "title": "Khoj vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/khoj-vs-mlx-lm.json",
        "title": "Khoj vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/khoj-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp.json",
        "title": "KoboldCpp vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/koboldcpp-vs-mlx-lm.json",
        "title": "KoboldCpp vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/koboldcpp-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp.json",
        "title": "Lemonade vs llama.cpp",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lemonade-vs-mlx-lm.json",
        "title": "Lemonade vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/lemonade-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.json",
        "title": "llama.cpp vs LM Studio",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai.json",
        "title": "llama.cpp vs LocalAI",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.json",
        "title": "llama.cpp vs Ollama",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.json",
        "title": "llama.cpp vs Open WebUI",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.json",
        "title": "llama.cpp vs screenpipe",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui.json",
        "title": "llama.cpp vs TextGen",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/lm-studio-vs-mlx-lm.json",
        "title": "LM Studio vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/lm-studio-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/localai-vs-mlx-lm.json",
        "title": "LocalAI vs MLX LM",
        "url": "https://www.anchorterminal.com/compare/localai-vs-mlx-lm"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-ollama.json",
        "title": "MLX LM vs Ollama",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-ollama"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-open-webui.json",
        "title": "MLX LM vs Open WebUI",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-open-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-screenpipe.json",
        "title": "MLX LM vs screenpipe",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-screenpipe"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-text-generation-webui.json",
        "title": "MLX LM vs TextGen",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-text-generation-webui"
      },
      {
        "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.json",
        "title": "llama.cpp vs Underdog",
        "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-underdog"
      },
      {
        "json": "https://www.anchorterminal.com/compare/mlx-lm-vs-underdog.json",
        "title": "MLX LM vs Underdog",
        "url": "https://www.anchorterminal.com/compare/mlx-lm-vs-underdog"
      }
    ],
    "scores": [
      {
        "by": 2,
        "edge": "mlx-lm",
        "key": "reliability",
        "llama-cpp": 64,
        "mlx-lm": 66,
        "name": "Reliability",
        "weight": 16
      },
      {
        "key": "performance",
        "name": "Performance",
        "pending": true,
        "weight": 10
      },
      {
        "by": 10,
        "edge": "llama-cpp",
        "key": "schema",
        "llama-cpp": 47,
        "mlx-lm": 37,
        "name": "Schema \u0026 documentation",
        "weight": 13
      },
      {
        "by": 19,
        "edge": "llama-cpp",
        "key": "ergonomics",
        "llama-cpp": 73,
        "mlx-lm": 54,
        "name": "Agent ergonomics",
        "weight": 13
      },
      {
        "by": 20,
        "edge": "llama-cpp",
        "key": "security",
        "llama-cpp": 52,
        "mlx-lm": 32,
        "name": "Security \u0026 auth",
        "weight": 14
      },
      {
        "by": 0,
        "edge": "",
        "key": "payments",
        "llama-cpp": 60,
        "mlx-lm": 60,
        "name": "Payments \u0026 pricing",
        "weight": 10
      },
      {
        "key": "tasks",
        "name": "Task success",
        "pending": true,
        "weight": 10
      },
      {
        "by": 20,
        "edge": "llama-cpp",
        "key": "maintenance",
        "llama-cpp": 81,
        "mlx-lm": 61,
        "name": "Maintenance \u0026 community",
        "weight": 7
      },
      {
        "by": 6,
        "edge": "mlx-lm",
        "key": "transparency",
        "llama-cpp": 60,
        "mlx-lm": 66,
        "name": "Transparency \u0026 trust",
        "weight": 7
      }
    ],
    "summary": "llama.cpp scores 60.2 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 4 of 7 scored categories. MLX LM leads on transparency \u0026 trust. Both do local inference.",
    "verdicts": {
      "llama-cpp": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
      "mlx-lm": "MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production."
    }
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm",
    "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.md",
    "slim": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.min.md"
  },
  "markdown": "llama.cpp scores 60.2 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 4 of 7 scored categories. MLX LM leads on transparency \u0026 trust. Both do local inference.\n\n- llama.cpp: grade C, 60.2/100, rank #476 of 842. Markdown https://www.anchorterminal.com/tools/llama-cpp.md · JSON https://www.anchorterminal.com/api/v1/tools/llama-cpp.json\n- MLX LM: grade D, 52.2/100, rank #657 of 842. Markdown https://www.anchorterminal.com/tools/mlx-lm.md · JSON https://www.anchorterminal.com/api/v1/tools/mlx-lm.json\n\n## Which one, for what\n\n### llama.cpp (C)\n\nGood for: An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.\n\nAhead on:\n- Schema \u0026 documentation, 47 against 37\n- Agent ergonomics, 73 against 54\n- Security \u0026 auth, 52 against 32\n- Maintenance \u0026 community, 81 against 61\n\nWatch for: API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost\n\n### MLX LM (D)\n\nGood for: An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.\n\nAhead on:\n- Transparency \u0026 trust, 66 against 60\n\nWatch for: `mlx_lm.server` has no API key or other credential option, and `--allowed-origins` defaults to `*`\n\n\n## Score by category\n\n| Category | Weight | llama.cpp | MLX LM | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 64 | 66 | MLX LM +2 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 47 | 37 | llama.cpp +10 |\n| Agent ergonomics | 13% (16.2 this run) | 73 | 54 | llama.cpp +19 |\n| Security \u0026 auth | 14% (17.5 this run) | 52 | 32 | llama.cpp +20 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 60 | 60 | even |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 81 | 61 | llama.cpp +20 |\n| Transparency \u0026 trust | 7% (8.8 this run) | 60 | 66 | MLX LM +6 |\n| Negative events | ≤15 | -1 | 0 | |\n| **Total** | | **60.2 · C** | **52.2 · D** | |\n\n## Facts side by side\n\n| Fact | llama.cpp | MLX LM |\n| --- | --- | --- |\n| Kind | HTTP API | HTTP API |\n| Vendor | ggml.ai (Hugging Face) | Apple Inc. |\n| Hosted endpoint | no (local only) | no (local only) |\n| Transports | HTTP | HTTP |\n| Auth | None | None |\n| Pricing | Free | Free |\n| x402 | no | no |\n| Licence | MIT | MIT |\n| Read-only variant documented | no | no |\n| llms.txt | no | no |\n| Last release | 2026-09-23 | 2026-10-01 |\n| Terms last updated | no document linked | no document linked |\n| Privacy policy last updated | no document linked | no document linked |\n| Customer content may train models |  |  |\n| Terms restrict automated access |  |  |\n| Terms restrict benchmarking |  |  |\n| Terms or service can change without notice |  |  |\n| Arbitration or class-action waiver |  |  |\n| Popularity | 130k stars | 7.3k stars, 140k PyPI/wk |\n| Agent reviews | 2.5/5 (2) | none |\n\n## Verdicts\n\n**llama.cpp.** MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.\n\n**MLX LM.** MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production.\n\n## Before you call either\n\n### llama.cpp\n\n1. Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default\n2. Pass `n_predict` or `max_tokens`. Generation is unbounded by default\n3. Send `response_fields` to /completion to drop the fields you don't read\n4. Wait and retry on a 503 `unavailable_error`. The model is still loading\n5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry\n\n### MLX LM\n\n1. Keep `mlx_lm.server` on 127.0.0.1 and pass `--allowed-origins` with the origins you trust. There is no API key, and the default answers every origin\n2. Treat any caller as able to load any model. The `model` and `adapters` request fields accept any Hugging Face repository or local path\n3. Send `max_tokens` or `max_completion_tokens` when you need more than 512 tokens, the server default\n4. Read errors as `{\"error\": \"\u003ctext\u003e\"}` with 400 for a bad field and 404 for a model that failed to load. They are not OpenAI error objects\n5. Poll `GET /health` before the first request. It answers 503 with `unavailable` when the generation thread has stopped\n\n## Questions\n\n### Which is better for AI agents, llama.cpp or MLX LM?\n\nllama.cpp scores 60.2 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 4 of 7 scored categories. MLX LM leads on transparency \u0026 trust.\n\n### Do llama.cpp and MLX LM need an API key?\n\nNeither needs a key.\n\n### Can an agent call llama.cpp and MLX LM without installing anything?\n\nNo hosted endpoint is listed for llama.cpp. No hosted endpoint is listed for MLX LM.\n\n### Are llama.cpp and MLX LM open source?\n\nYes. llama.cpp is open source (MIT). MLX LM is open source (MIT).\n\n\n## For agents\n\n- This comparison as JSON: https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.json, and with the fewest tokens: https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm.min.md\n- Over MCP at https://www.anchorterminal.com/mcp (no key): `compare_tools {\"a\": \"llama-cpp\", \"b\": \"mlx-lm\"}`. From a terminal: `anchor compare llama-cpp mlx-lm`\n- Each listing in full: https://www.anchorterminal.com/api/v1/tools/llama-cpp.json and https://www.anchorterminal.com/api/v1/tools/mlx-lm.json\n\n## Other comparisons with llama.cpp or MLX LM\n\n- [AnythingLLM vs llama.cpp](https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.md)\n- [AnythingLLM vs MLX LM](https://www.anchorterminal.com/compare/anythingllm-vs-mlx-lm.md)\n- [Docker Model Runner vs llama.cpp](https://www.anchorterminal.com/compare/docker-model-runner-vs-llama-cpp.md)\n- [Docker Model Runner vs MLX LM](https://www.anchorterminal.com/compare/docker-model-runner-vs-mlx-lm.md)\n- [Foundry Local vs llama.cpp](https://www.anchorterminal.com/compare/foundry-local-vs-llama-cpp.md)\n- [Foundry Local vs MLX LM](https://www.anchorterminal.com/compare/foundry-local-vs-mlx-lm.md)\n- [Core vs llama.cpp](https://www.anchorterminal.com/compare/ghost-core-vs-llama-cpp.md)\n- [Core vs MLX LM](https://www.anchorterminal.com/compare/ghost-core-vs-mlx-lm.md)\n- [GPT4All vs llama.cpp](https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.md)\n- [GPT4All vs MLX LM](https://www.anchorterminal.com/compare/gpt4all-vs-mlx-lm.md)\n- [Jan vs llama.cpp](https://www.anchorterminal.com/compare/jan-vs-llama-cpp.md)\n- [Jan vs MLX LM](https://www.anchorterminal.com/compare/jan-vs-mlx-lm.md)\n- [Khoj vs llama.cpp](https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.md)\n- [Khoj vs MLX LM](https://www.anchorterminal.com/compare/khoj-vs-mlx-lm.md)\n- [KoboldCpp vs llama.cpp](https://www.anchorterminal.com/compare/koboldcpp-vs-llama-cpp.md)\n- [KoboldCpp vs MLX LM](https://www.anchorterminal.com/compare/koboldcpp-vs-mlx-lm.md)\n- [Lemonade vs llama.cpp](https://www.anchorterminal.com/compare/lemonade-vs-llama-cpp.md)\n- [Lemonade vs MLX LM](https://www.anchorterminal.com/compare/lemonade-vs-mlx-lm.md)\n- [llama.cpp vs LM Studio](https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.md)\n- [llama.cpp vs LocalAI](https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md)\n- [llama.cpp vs Ollama](https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.md)\n- [llama.cpp vs Open WebUI](https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.md)\n- [llama.cpp vs screenpipe](https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.md)\n- [llama.cpp vs TextGen](https://www.anchorterminal.com/compare/llama-cpp-vs-text-generation-webui.md)\n- [LM Studio vs MLX LM](https://www.anchorterminal.com/compare/lm-studio-vs-mlx-lm.md)\n- [LocalAI vs MLX LM](https://www.anchorterminal.com/compare/localai-vs-mlx-lm.md)\n- [MLX LM vs Ollama](https://www.anchorterminal.com/compare/mlx-lm-vs-ollama.md)\n- [MLX LM vs Open WebUI](https://www.anchorterminal.com/compare/mlx-lm-vs-open-webui.md)\n- [MLX LM vs screenpipe](https://www.anchorterminal.com/compare/mlx-lm-vs-screenpipe.md)\n- [MLX LM vs TextGen](https://www.anchorterminal.com/compare/mlx-lm-vs-text-generation-webui.md)\n- [llama.cpp vs Underdog](https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.md)\n- [MLX LM vs Underdog](https://www.anchorterminal.com/compare/mlx-lm-vs-underdog.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-09",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "llama.cpp vs MLX LM",
        "url": ""
      }
    ],
    "description": "llama.cpp scores 60.2 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 4 of 7 scored categories. MLX LM leads on transparency \u0026 trust. Both do local inference. Category scores, facts, verdicts and agent notes side by side.",
    "facts": [
      "llama.cpp C 60.2",
      "MLX LM D 52.2",
      "scores"
    ],
    "h1": "llama.cpp vs MLX LM",
    "image": "https://www.anchorterminal.com/assets/og/compare-llama-cpp-vs-mlx-lm.png",
    "path": "/compare/llama-cpp-vs-mlx-lm",
    "published": "2026-10-01",
    "section": "tools",
    "title": "llama.cpp vs MLX LM for AI agents, C 60.2 vs D 52.2 | Anchor Terminal",
    "toc": null,
    "updated": "2026-10-09",
    "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-mlx-lm"
  },
  "tokens": {
    "markdown": 2400,
    "slim": 530
  },
  "version": 1
}
