{
  "data": {
    "a": {
      "slug": "llama-cpp",
      "name": "llama.cpp",
      "vendor": "ggml.ai (Hugging Face)",
      "vendorUrl": "https://llama.app",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source C/C++ engine for running GGUF models locally, with a web interface and compatible model APIs.",
      "url": "https://www.anchorterminal.com/tools/llama-cpp",
      "markdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/llama-cpp.json",
      "repo": "https://github.com/ggml-org/llama.cpp",
      "license": "MIT",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "ghcr.io/ggml-org/llama.cpp"
        },
        {
          "registry": "pypi",
          "name": "gguf"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one key or a comma-separated list) or `--api-key-file` (one key a line) turns on a check for every route but /health and the web UI's files, with the key sent as `Authorization: Bearer` or `X-Api-Key`, never in the query string. Keys have no scopes and change only with a restart. TLS is built in with `--ssl-key-file` and `--ssl-cert-file`. The server binds 127.0.0.1:8080 by default, and CORS reflects any Origin with credentials allowed unless built-in tools, MCP servers or `--agent` are on, when it narrows to localhost (https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).",
      "pricing": "free",
      "pricingNotes": "Free under MIT, with no account, key or card. Nothing is sold. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 130200,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "no-telemetry"
      ],
      "lastRelease": "2026-09-23",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 60.2,
        "grade": "C",
        "agentReady": false,
        "rank": 253,
        "ranked": true,
        "rankOf": 452,
        "categoryRank": 3,
        "methodology": "0.3",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 60
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -1,
        "negativeNotes": [
          "2026-03-26. GHSA-j8rj-fmpv-wcxw (CVE-2026-34159, 9.8 at NVD), unauthenticated code execution through a GRAPH_COMPUTE bypass in the RPC backend, the most serious of four advisories published between January and March 2026 (the others a llama-server out-of-bounds write through a negative `n_discard` and two GGUF integer overflows). All were fixed in named builds and published as advisories, SECURITY.md says not to expose the RPC server or llama-server to untrusted networks, and the newest is more than six months old, -1. https://github.com/ggml-org/llama.cpp/security/advisories/GHSA-j8rj-fmpv-wcxw; https://github.com/ggml-org/llama.cpp/security"
        ],
        "verdict": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
        "strengths": [
          "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads",
          "OpenAI chat completions, responses and embeddings, Anthropic messages, reranking and /v1/systemone from one server",
          "`response_fields`, `json_schema` and `grammar` control the size and shape of output, and errors carry an OpenAI-style type and code",
          "1,005 nightly builds and eight semver releases in 90 days, with 37 workflows running on every push to master",
          "Ten published GitHub advisories with CVEs and fixed builds, and SECURITY.md guidance on untrusted models and inputs"
        ],
        "weaknesses": [
          "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost",
          "No OpenAPI file of its own, and the REST API changelog stops at b4599",
          "Private security disclosure disabled since 1 June 2026, with fixes asked for as public pull requests",
          "Pre-1.0 (0.5.0), and semver releases are bare tags with no notes",
          "No official client library, and `n_predict` defaults to unlimited"
        ],
        "agentNotes": [
          "Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default",
          "Pass `n_predict` or `max_tokens`. Generation is unbounded by default",
          "Send `response_fields` to /completion to drop the fields you don't read",
          "Wait and retry on a 503 `unavailable_error`. The model is still loading",
          "Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 2.5,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.3",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 60.2
          }
        ],
        "editorialScores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 66
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "curl -LsSf https://llama.app/install.sh | sh   # or: brew install llama.cpp; winget install llama.cpp\nllama serve -hf ggml-org/Qwen3.5-0.8B-GGUF   # listens on 127.0.0.1:8080",
        "http": "curl --request POST \\\n    --url http://localhost:8080/completion \\\n    --header \"Content-Type: application/json\" \\\n    --data '{\"prompt\": \"Building a website can be done in 10 simple steps:\",\"n_predict\": 128}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/llama-cpp"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "ggml.ai, part of Hugging Face since 2026",
        "domain": "llama.app",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/ggml-org/llama.cpp/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "The repository's About link is llama.app, which says it's by the llama.cpp team and Hugging Face and links no terms, privacy or security page. ggml.ai says the company was acquired by Hugging Face in 2026 and names no address.",
          "The `LICENSE` file reads Copyright (c) 2023-2026 The ggml authors.",
          "llama.app/.well-known/security.txt and llama.app/llms.txt return 404. SECURITY.md points to GitHub private advisories while saying private disclosure is disabled.",
          "There's no shared hosted endpoint. The server runs on the owner's machine."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/llama-cpp.json",
      "live": {
        "slug": "llama-cpp",
        "versions": [
          {
            "registry": "github",
            "name": "ggml-org/llama.cpp",
            "version": "v0.5.0",
            "released": "2026-09-23",
            "seenAt": "2026-10-04T16:31:53.040249174Z"
          },
          {
            "registry": "pypi",
            "name": "gguf",
            "version": "0.19.0",
            "released": "2026-05-06",
            "seenAt": "2026-10-04T16:31:52.933067593Z"
          }
        ],
        "githubStars": 130286,
        "securityTxt": {
          "url": "https://llama.app/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-04T15:16:00.86400098Z"
        },
        "domain": {
          "domain": "llama.app",
          "registered": "2018-07-18",
          "source": "https://pubapi.registry.google/rdap/domain/llama.app",
          "checkedAt": "2026-10-04T13:04:03.05886804Z"
        },
        "updatedAt": "2026-10-04T16:31:53.040249174Z"
      }
    },
    "b": {
      "slug": "localai",
      "name": "LocalAI",
      "vendor": "Ettore Di Giacinto and the LocalAI team",
      "vendorUrl": "https://localai.io",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source engine in Go, MIT licensed, that runs models on the owner's hardware behind OpenAI-, Anthropic-, Ollama- and ElevenLabs-compatible APIs on port 8080.",
      "url": "https://www.anchorterminal.com/tools/localai",
      "markdownUrl": "https://www.anchorterminal.com/tools/localai.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/localai.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/localai.json",
      "repo": "https://github.com/mudler/LocalAI",
      "license": "MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence",
      "transports": [
        "http",
        "stdio"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "docker.io/localai/localai"
        }
      ],
      "auth": "mixed",
      "authNotes": "Off by default. With no keys and no user accounts configured, every request is accepted, and the server refuses to start on a public address in that state unless `--allow-insecure-public-bind` is set. `LOCALAI_API_KEY` sets shared keys with full admin rights. `LOCALAI_AUTH=true` turns on user accounts (local, GitHub OAuth or OIDC), and each user creates revocable keys stored as HMAC-SHA256, with an optional expiry in the source, carrying the user's role (admin or user) and per-model and per-feature permissions. Keys go in `Authorization: Bearer`, `x-api-key`, `xi-api-key` or a `token` cookie.",
      "pricing": "free",
      "pricingNotes": "Free and MIT with nothing to buy. You run it on your own hardware. The project takes sponsorship through GitHub Sponsors.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": 42,
      "popularity": {
        "githubStars": 47800,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://localai.io/basics/getting_started/",
      "openapi": "https://raw.githubusercontent.com/mudler/LocalAI/master/swagger/swagger.json",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "agent.mcp-client",
        "embed.text",
        "rerank",
        "speech.stt",
        "speech.tts",
        "voice.speech-to-speech",
        "image.generate",
        "video.generate",
        "guard.pii",
        "finetune.sft",
        "db.vector"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "openapi",
        "mcp",
        "go",
        "docker",
        "streaming",
        "open-weights",
        "no-telemetry"
      ],
      "lastRelease": "2026-10-02",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 68,
        "grade": "B",
        "agentReady": false,
        "rank": 133,
        "ranked": true,
        "rankOf": 452,
        "categoryRank": 1,
        "methodology": "0.3",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 71,
          "maintenance": 80,
          "payments": 60,
          "reliability": 84,
          "schema": 81,
          "security": 62,
          "transparency": 47
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -3,
        "negativeNotes": [
          "2026-07-07. CVE-2026-59707 (8.6 at NVD under CVSS 3.1, published by VulnCheck), an unauthenticated server-side request forgery through POST /models/apply in v4.3.1 and earlier, reported in issue #10665. The code now refuses private, loopback and metadata addresses in gallery config fetches, with a comment citing the issue, from v4.8.0 at the latest. The project published no GitHub advisory, the fix commit NVD and VulnCheck name (f9b968e) is an unrelated docs change, and SECURITY.md still lists 3.x as the supported series. Fixed, documented only by a third party, -3. https://nvd.nist.gov/vuln/detail/CVE-2026-59707; https://github.com/mudler/LocalAI/issues/10665"
        ],
        "verdict": "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.",
        "strengths": [
          "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app",
          "OpenAI, Anthropic, Open Responses, Ollama and ElevenLabs-compatible endpoints, with a Swagger 2.0 file of 133 operations served by every instance",
          "429 and 503 responses carry Retry-After, and errors come in the calling client's own envelope",
          "Optional user accounts with hashed, revocable keys, per-model and per-feature permissions and per-user quotas",
          "v4.11.0 on 2 October 2026, ten releases in 90 days, and the Tests workflow passing on master"
        ],
        "weaknesses": [
          "No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin",
          "CVE-2026-59707, an unauthenticated SSRF in v4.3.1 and earlier, published by VulnCheck in July 2026 with no advisory from the project",
          "SECURITY.md still names 3.x as the supported series, and there's no security.txt or privacy policy",
          "The MCP admin server registers 42 tools against the 19 its docs list, with no annotations, and its writes are held back only by a prompt",
          "No breaking-change section in the release notes, and the unsigned macOS DMG needs its quarantine flag removed by hand"
        ],
        "agentNotes": [
          "Send `Authorization: Bearer \u003ckey\u003e` when the operator has set keys. A 401 means the instance has auth on",
          "Read /.well-known/localai.json and /api/instructions first. Both answer without a key and list what this instance can do",
          "Back off on 429 and 503 for the Retry-After seconds. A 503 can mean the model is still loading",
          "Start `local-ai mcp-server` with `--read-only` unless the task is to install or delete models",
          "Take model names from /v1/models. Each instance names its own"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 3,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "B",
            "methodology": "0.3",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 68
          }
        ],
        "editorialScores": {
          "ergonomics": 71,
          "maintenance": 80,
          "payments": 60,
          "reliability": 84,
          "schema": 81,
          "security": 62,
          "transparency": 67
        },
        "provenanceScore": 27
      },
      "connect": {
        "install": "docker run -ti --name local-ai -p 8080:8080 localai/localai:latest",
        "http": "curl http://localhost:8080/v1/chat/completions -H \"Content-Type: application/json\" -d '{\n  \"model\": \"qwen3-4b\",\n  \"messages\": [{\"role\": \"user\", \"content\": \"Hello!\"}]\n}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/localai"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "",
        "domain": "localai.io",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/mudler/LocalAI/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "No company is named. The `LICENSE` copyright line reads Ettore Di Giacinto, and the README names him as project lead with Richard Palethorpe as maintainer.",
          "We found no terms or privacy page in the docs site's source, and localai.io/.well-known/security.txt and localai.io/llms.txt return 404.",
          "There's no hosted endpoint. Each instance answers on the operator's own host, by default port 8080."
        ],
        "score": 27
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/localai.json",
      "live": {
        "slug": "localai",
        "versions": [
          {
            "registry": "github",
            "name": "mudler/LocalAI",
            "version": "v4.11.0",
            "released": "2026-10-02",
            "seenAt": "2026-10-04T16:32:02.572449913Z"
          }
        ],
        "githubStars": 49385,
        "securityTxt": {
          "url": "https://localai.io/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-04T15:15:42.873292356Z"
        },
        "domain": {
          "domain": "localai.io",
          "checkedAt": "2026-10-04T13:07:02.946116654Z"
        },
        "updatedAt": "2026-10-04T16:32:02.572449913Z"
      }
    },
    "summary": "LocalAI has a score of 68 (B) against llama.cpp's 60.2 (C). Both do local inference. The largest gap is schema \u0026 documentation, 34 points."
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai",
    "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md",
    "slim": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai.min.md"
  },
  "markdown": "LocalAI has a score of 68 (B) against llama.cpp's 60.2 (C). Both do local inference. The largest gap is schema \u0026 documentation, 34 points.\n\n- llama.cpp: grade C, 60.2/100, rank #253 of 452. Markdown https://www.anchorterminal.com/tools/llama-cpp.md · JSON https://www.anchorterminal.com/api/v1/tools/llama-cpp.json\n- LocalAI: grade B, 68/100, rank #133 of 452. Markdown https://www.anchorterminal.com/tools/localai.md · JSON https://www.anchorterminal.com/api/v1/tools/localai.json\n\n## Which one, for what\n\nPick llama.cpp for transparency \u0026 trust (+13).\n\nPick LocalAI for reliability (+20), schema \u0026 documentation (+34), security \u0026 auth (+10).\n\n## Score by category\n\n| Category | Weight | llama.cpp | LocalAI | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 64 | 84 | LocalAI +20 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 47 | 81 | LocalAI +34 |\n| Agent ergonomics | 13% (16.2 this run) | 73 | 71 | llama.cpp +2 |\n| Security \u0026 auth | 14% (17.5 this run) | 52 | 62 | LocalAI +10 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 60 | 60 | even |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 81 | 80 | llama.cpp +1 |\n| Transparency \u0026 trust | 7% (8.8 this run) | 60 | 47 | llama.cpp +13 |\n| Negative events | ≤15 | -1 | -3 | |\n| **Total** | | **60.2 · C** | **68 · B** | |\n\n## Facts side by side\n\n| Fact | llama.cpp | LocalAI |\n| --- | --- | --- |\n| Kind | HTTP API | HTTP API |\n| Vendor | ggml.ai (Hugging Face) | Ettore Di Giacinto and the LocalAI team |\n| Hosted endpoint | no (local only) | no (local only) |\n| Transports | HTTP | HTTP, stdio |\n| Auth | None | OAuth or key |\n| Pricing | Free | Free |\n| x402 | no | no |\n| Licence | MIT | MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence |\n| Tools exposed | none | 42 |\n| Context cost (tools/list) | n/a | n/a |\n| p95 latency | not measured yet | not measured yet |\n| Availability (30d) | not measured yet | not measured yet |\n| Read-only variant documented | no | yes |\n| llms.txt | no | no |\n| MCP registry | not listed | not listed |\n| Last release | 2026-09-23 | 2026-10-02 |\n| Popularity | 130k stars | 48k stars |\n| Agent reviews | 2.5/5 (2) | 3/5 (2) |\n\n## Verdicts\n\n**llama.cpp.** MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.\n\n**LocalAI.** MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.\n\n## Before you call either\n\n### llama.cpp\n\n1. Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default\n2. Pass `n_predict` or `max_tokens`. Generation is unbounded by default\n3. Send `response_fields` to /completion to drop the fields you don't read\n4. Wait and retry on a 503 `unavailable_error`. The model is still loading\n5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry\n\n### LocalAI\n\n1. Send `Authorization: Bearer \u003ckey\u003e` when the operator has set keys. A 401 means the instance has auth on\n2. Read /.well-known/localai.json and /api/instructions first. Both answer without a key and list what this instance can do\n3. Back off on 429 and 503 for the Retry-After seconds. A 503 can mean the model is still loading\n4. Start `local-ai mcp-server` with `--read-only` unless the task is to install or delete models\n5. Take model names from /v1/models. Each instance names its own\n\n## Other comparisons with llama.cpp or LocalAI\n\n- [AnythingLLM vs llama.cpp](https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.md)\n- [AnythingLLM vs LocalAI](https://www.anchorterminal.com/compare/anythingllm-vs-localai.md)\n- [GPT4All vs llama.cpp](https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.md)\n- [GPT4All vs LocalAI](https://www.anchorterminal.com/compare/gpt4all-vs-localai.md)\n- [Jan vs llama.cpp](https://www.anchorterminal.com/compare/jan-vs-llama-cpp.md)\n- [Jan vs LocalAI](https://www.anchorterminal.com/compare/jan-vs-localai.md)\n- [Khoj vs llama.cpp](https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.md)\n- [Khoj vs LocalAI](https://www.anchorterminal.com/compare/khoj-vs-localai.md)\n- [llama.cpp vs LM Studio](https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.md)\n- [llama.cpp vs Ollama](https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.md)\n- [llama.cpp vs Open WebUI](https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.md)\n- [llama.cpp vs screenpipe](https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.md)\n- [LM Studio vs LocalAI](https://www.anchorterminal.com/compare/lm-studio-vs-localai.md)\n- [LocalAI vs Ollama](https://www.anchorterminal.com/compare/localai-vs-ollama.md)\n- [LocalAI vs Open WebUI](https://www.anchorterminal.com/compare/localai-vs-open-webui.md)\n- [LocalAI vs screenpipe](https://www.anchorterminal.com/compare/localai-vs-screenpipe.md)\n- [llama.cpp vs Underdog](https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.md)\n- [LocalAI vs Underdog](https://www.anchorterminal.com/compare/localai-vs-underdog.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-05",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.3",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "llama.cpp vs LocalAI",
        "url": ""
      }
    ],
    "description": "LocalAI has a score of 68 (B) against llama.cpp's 60.2 (C). Both do local inference. The largest gap is schema \u0026 documentation, 34 points. Category scores, facts, verdicts and agent notes side by side.",
    "facts": [
      "llama.cpp C 60.2",
      "LocalAI B 68",
      "scores"
    ],
    "h1": "llama.cpp vs LocalAI",
    "image": "https://www.anchorterminal.com/assets/og/compare-llama-cpp-vs-localai.png",
    "path": "/compare/llama-cpp-vs-localai",
    "published": "2026-10-01",
    "section": "tools",
    "title": "llama.cpp vs LocalAI for AI agents, C 60.2 vs B 68 | Anchor Terminal",
    "toc": null,
    "updated": "2026-10-05",
    "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-localai"
  },
  "tokens": {
    "markdown": 1550,
    "slim": 330
  },
  "version": 1
}
