{
  "data": {
    "a": {
      "slug": "llama-cpp",
      "name": "llama.cpp",
      "vendor": "ggml.ai (Hugging Face)",
      "vendorUrl": "https://llama.app",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source C/C++ engine for running GGUF models locally, with a web interface and compatible model APIs.",
      "url": "https://www.anchorterminal.com/tools/llama-cpp",
      "markdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/llama-cpp.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/llama-cpp.json",
      "repo": "https://github.com/ggml-org/llama.cpp",
      "license": "MIT",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "ghcr.io/ggml-org/llama.cpp"
        },
        {
          "registry": "pypi",
          "name": "gguf"
        }
      ],
      "auth": "none",
      "authNotes": "No credential by default. `--api-key` (one key or a comma-separated list) or `--api-key-file` (one key a line) turns on a check for every route but /health and the web UI's files, with the key sent as `Authorization: Bearer` or `X-Api-Key`, never in the query string. Keys have no scopes and change only with a restart. TLS is built in with `--ssl-key-file` and `--ssl-cert-file`. The server binds 127.0.0.1:8080 by default, and CORS reflects any Origin with credentials allowed unless built-in tools, MCP servers or `--agent` are on, when it narrows to localhost (https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).",
      "pricing": "free",
      "pricingNotes": "Free under MIT, with no account, key or card. Nothing is sold. You pay for your own hardware and electricity.",
      "priceSummary": "Free · OSS",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 130200,
        "npmWeekly": null,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "embed.text",
        "rerank",
        "inference.decision",
        "agent.mcp-client"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "free",
        "no-card",
        "openai-compatible",
        "docker",
        "pre-1.0",
        "no-telemetry"
      ],
      "lastRelease": "2026-09-23",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 60.2,
        "grade": "C",
        "agentReady": false,
        "rank": 253,
        "ranked": true,
        "rankOf": 452,
        "categoryRank": 3,
        "methodology": "0.3",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 60
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -1,
        "negativeNotes": [
          "2026-03-26. GHSA-j8rj-fmpv-wcxw (CVE-2026-34159, 9.8 at NVD), unauthenticated code execution through a GRAPH_COMPUTE bypass in the RPC backend, the most serious of four advisories published between January and March 2026 (the others a llama-server out-of-bounds write through a negative `n_discard` and two GGUF integer overflows). All were fixed in named builds and published as advisories, SECURITY.md says not to expose the RPC server or llama-server to untrusted networks, and the newest is more than six months old, -1. https://github.com/ggml-org/llama.cpp/security/advisories/GHSA-j8rj-fmpv-wcxw; https://github.com/ggml-org/llama.cpp/security"
        ],
        "verdict": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
        "strengths": [
          "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads",
          "OpenAI chat completions, responses and embeddings, Anthropic messages, reranking and /v1/systemone from one server",
          "`response_fields`, `json_schema` and `grammar` control the size and shape of output, and errors carry an OpenAI-style type and code",
          "1,005 nightly builds and eight semver releases in 90 days, with 37 workflows running on every push to master",
          "Ten published GitHub advisories with CVEs and fixed builds, and SECURITY.md guidance on untrusted models and inputs"
        ],
        "weaknesses": [
          "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost",
          "No OpenAPI file of its own, and the REST API changelog stops at b4599",
          "Private security disclosure disabled since 1 June 2026, with fixes asked for as public pull requests",
          "Pre-1.0 (0.5.0), and semver releases are bare tags with no notes",
          "No official client library, and `n_predict` defaults to unlimited"
        ],
        "agentNotes": [
          "Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default",
          "Pass `n_predict` or `max_tokens`. Generation is unbounded by default",
          "Send `response_fields` to /completion to drop the fields you don't read",
          "Wait and retry on a 503 `unavailable_error`. The model is still loading",
          "Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 2.5,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.3",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 60.2
          }
        ],
        "editorialScores": {
          "ergonomics": 73,
          "maintenance": 81,
          "payments": 60,
          "reliability": 64,
          "schema": 47,
          "security": 52,
          "transparency": 66
        },
        "provenanceScore": 53
      },
      "connect": {
        "install": "curl -LsSf https://llama.app/install.sh | sh   # or: brew install llama.cpp; winget install llama.cpp\nllama serve -hf ggml-org/Qwen3.5-0.8B-GGUF   # listens on 127.0.0.1:8080",
        "http": "curl --request POST \\\n    --url http://localhost:8080/completion \\\n    --header \"Content-Type: application/json\" \\\n    --data '{\"prompt\": \"Building a website can be done in 10 simple steps:\",\"n_predict\": 128}'"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/llama-cpp"
      },
      "area": "models",
      "provenance": {
        "legalEntity": "ggml.ai, part of Hugging Face since 2026",
        "domain": "llama.app",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "",
        "privacy": "",
        "statusPage": "",
        "changelog": "https://github.com/ggml-org/llama.cpp/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "The repository's About link is llama.app, which says it's by the llama.cpp team and Hugging Face and links no terms, privacy or security page. ggml.ai says the company was acquired by Hugging Face in 2026 and names no address.",
          "The `LICENSE` file reads Copyright (c) 2023-2026 The ggml authors.",
          "llama.app/.well-known/security.txt and llama.app/llms.txt return 404. SECURITY.md points to GitHub private advisories while saying private disclosure is disabled.",
          "There's no shared hosted endpoint. The server runs on the owner's machine."
        ],
        "score": 53
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/llama-cpp.json",
      "live": {
        "slug": "llama-cpp",
        "versions": [
          {
            "registry": "github",
            "name": "ggml-org/llama.cpp",
            "version": "v0.5.0",
            "released": "2026-09-23",
            "seenAt": "2026-10-04T16:31:53.040249174Z"
          },
          {
            "registry": "pypi",
            "name": "gguf",
            "version": "0.19.0",
            "released": "2026-05-06",
            "seenAt": "2026-10-04T16:31:52.933067593Z"
          }
        ],
        "githubStars": 130286,
        "securityTxt": {
          "url": "https://llama.app/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-04T15:16:00.86400098Z"
        },
        "domain": {
          "domain": "llama.app",
          "registered": "2018-07-18",
          "source": "https://pubapi.registry.google/rdap/domain/llama.app",
          "checkedAt": "2026-10-04T13:04:03.05886804Z"
        },
        "updatedAt": "2026-10-04T16:31:53.040249174Z"
      }
    },
    "b": {
      "slug": "ollama",
      "name": "Ollama",
      "vendor": "Ollama Inc.",
      "vendorUrl": "https://ollama.com",
      "kind": "http-api",
      "category": "local-ai",
      "summary": "Open-source model runner for macOS, Windows and Linux, with a local API and a library of downloadable models.",
      "url": "https://www.anchorterminal.com/tools/ollama",
      "markdownUrl": "https://www.anchorterminal.com/tools/ollama.md",
      "slimMarkdownUrl": "https://www.anchorterminal.com/tools/ollama.min.md",
      "jsonUrl": "https://www.anchorterminal.com/api/v1/tools/ollama.json",
      "repo": "https://github.com/ollama/ollama",
      "license": "MIT (server, CLI and desktop app). Ollama Cloud is a closed service under the ollama.com terms, and each model carries its own licence",
      "transports": [
        "http"
      ],
      "packages": [
        {
          "registry": "oci",
          "name": "docker.io/ollama/ollama"
        },
        {
          "registry": "pypi",
          "name": "ollama"
        },
        {
          "registry": "npm",
          "name": "ollama"
        }
      ],
      "auth": "none",
      "authNotes": "The local API at http://localhost:11434 takes no credential. It binds 127.0.0.1, answers a foreign Host header with 403 while bound to loopback, and allows cross-origin calls from 127.0.0.1 and 0.0.0.0 unless `OLLAMA_ORIGINS` adds more. Anything that reaches the port can generate, pull, push, create, copy and delete models. Cloud models through the local server need `ollama signin`, which signs requests with the install's own key. Direct calls to https://ollama.com/api and /v1 need a Bearer API key from ollama.com/settings/keys, which doesn't expire and has no scopes, and is revoked from the same page (https://github.com/ollama/ollama/blob/main/docs/api/authentication.mdx).",
      "pricing": "freemium",
      "pricingNotes": "The server, CLI and desktop app are free under MIT with no account. Ollama Cloud has five plans on ollama.com/pricing. Free ($0, starter usage credits, starter models, 1 concurrent request), Pro ($20 a month or $200 a year, $60 of usage credits a month, 3 concurrent requests), Max ($100 a month, $300 of credits, 10 concurrent requests), Team ($500 a month, $1,000 of shared credits, unlimited users) and Enterprise (custom). Usage is priced per model by the token, and the page doesn't say whether the Free plan needs a card (checked 2026-10-03).",
      "priceSummary": "$20 / mo",
      "where": "local",
      "x402": {
        "level": "no",
        "evidence": "No x402, MPP or L402 in the docs, the pricing page or the source (checked 2026-10-03).",
        "endpoints": []
      },
      "toolCount": null,
      "popularity": {
        "githubStars": 181200,
        "npmWeekly": 871543,
        "pypiWeekly": null,
        "asOf": "2026-10-03"
      },
      "docsUrl": "https://docs.ollama.com",
      "llmsTxt": "https://docs.ollama.com/llms.txt",
      "openapi": "https://raw.githubusercontent.com/ollama/ollama/main/docs/openapi.yaml",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "inference.llm",
        "embed.text",
        "inference.decision",
        "web.search",
        "web.fetch"
      ],
      "tags": [
        "open-source",
        "local",
        "self-hosted",
        "hosted",
        "freemium",
        "no-card",
        "openai-compatible",
        "openapi",
        "llms-txt",
        "docker",
        "go",
        "python",
        "typescript",
        "pre-1.0",
        "no-auth"
      ],
      "lastRelease": "2026-10-01",
      "graded": true,
      "anchor": {
        "graded": true,
        "score": 56.6,
        "grade": "C",
        "agentReady": false,
        "rank": 302,
        "ranked": true,
        "rankOf": 452,
        "categoryRank": 5,
        "methodology": "0.3",
        "run": "2026-10-01",
        "scores": {
          "ergonomics": 75,
          "maintenance": 81,
          "payments": 60,
          "reliability": 53,
          "schema": 79,
          "security": 28,
          "transparency": 63
        },
        "pending": [
          "performance",
          "tasks"
        ],
        "assessment": {
          "confidence": "medium",
          "date": "2026-10-03"
        },
        "negative": -4,
        "negativeNotes": [
          "2026-04-29. CERT Polska published CVE-2026-42248 and CVE-2026-42249 (9.8 each). The Windows app accepted downloaded updates without a signature check and took the file name from the server's response, and it installs updates silently, so whoever could answer the update request could run code on the machine. CERT Polska tested 0.12.10 to 0.17.5, and the Windows check stayed a stub returning success until v0.23.3 on 12 May 2026, whose notes list the fix only as `app: harden update flows`. CERT Polska says the maintainers didn't respond with details or the vulnerable range, and Ollama published no advisory. Fixed, but not disclosed by the vendor, -4. https://cert.pl/en/posts/2026/04/CVE-2026-42248/; https://github.com/ollama/ollama/releases/tag/v0.23.3"
        ],
        "verdict": "An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages. No credential on the local API, and any caller that reaches it can pull, push, create and delete models.",
        "strengths": [
          "An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages",
          "Native, OpenAI-compatible and Anthropic-compatible routes on one local port, with `ollama launch` for Claude Code, Codex and OpenCode",
          "28 releases in the 90 days to 3 October 2026, and official Python and JavaScript libraries released on 28 September",
          "Local prompts stay on the machine, and `OLLAMA_NO_CLOUD=1` turns off cloud models and web search",
          "Binds 127.0.0.1 by default and refuses foreign Host headers while bound to loopback"
        ],
        "weaknesses": [
          "No credential on the local API, and any caller that reaches it can pull, push, create and delete models",
          "No GitHub security advisory, against 12 CVEs on NVD since October 2025",
          "The Windows updater installed unsigned files until v0.23.3 on 12 May 2026, fixed under a release note that didn't mention security",
          "The desktop app checks ollama.com every hour with a signed request, even with automatic updates off, and no documented way to stop it",
          "A default context of 4k tokens below 24 GiB of VRAM, where the docs say agents need 64,000"
        ],
        "agentNotes": [
          "Send `\"stream\": false` for one JSON body. The native routes stream NDJSON by default",
          "Set `OLLAMA_CONTEXT_LENGTH=64000` or `options.num_ctx` before agent work. The default is 4k below 24 GiB of VRAM",
          "Back off on a 503. It means the queue (512 by default) is full",
          "Put an authenticating proxy in front before binding past 127.0.0.1. The server checks no credential",
          "Expect model names with a `cloud` tag to run on Ollama's servers. They need `ollama signin` and fail with `OLLAMA_NO_CLOUD=1`"
        ],
        "metrics": {
          "kind": "local",
          "measured": false
        },
        "reviewCount": 2,
        "avgRating": 2.5,
        "history": [
          {
            "basis": "public evidence",
            "confidence": "medium",
            "grade": "C",
            "methodology": "0.3",
            "pending": [
              "performance",
              "tasks"
            ],
            "run": "2026-10-01",
            "runLabel": "October 2026 research run",
            "score": 56.6
          }
        ],
        "editorialScores": {
          "ergonomics": 75,
          "maintenance": 81,
          "payments": 60,
          "reliability": 53,
          "schema": 79,
          "security": 28,
          "transparency": 66
        },
        "provenanceScore": 59
      },
      "connect": {
        "install": "curl -fsSL https://ollama.com/install.sh | sh   # macOS and Linux; Windows: irm https://ollama.com/install.ps1 | iex\nollama pull gemma4:e2b",
        "http": "curl http://localhost:11434/api/chat \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"model\": \"gemma4:e2b\",\n    \"messages\": [{\"role\": \"user\", \"content\": \"Say hello in one sentence.\"}],\n    \"stream\": false\n  }'",
        "claudeCode": "ollama launch claude   # or: ANTHROPIC_AUTH_TOKEN=ollama ANTHROPIC_API_KEY=\"\" ANTHROPIC_BASE_URL=http://localhost:11434 claude --model qwen3.5"
      },
      "letme": {
        "capability": "https://letme.dev/inference.local",
        "tool": "https://letme.dev/ollama"
      },
      "area": "models",
      "unitPrices": [
        {
          "item": "Ollama Cloud Pro",
          "unit": "month",
          "usd": 20,
          "note": "$60 of usage credits a month, 3 concurrent requests. $200 a year"
        },
        {
          "item": "Ollama Cloud Max",
          "unit": "month",
          "usd": 100,
          "note": "$300 of usage credits a month, 10 concurrent requests"
        },
        {
          "item": "Ollama Cloud Team",
          "unit": "month",
          "usd": 500,
          "note": "$1,000 of shared usage credits a month, unlimited users, 10 concurrent requests"
        }
      ],
      "provenance": {
        "legalEntity": "Ollama Inc.",
        "domain": "ollama.com",
        "domainRegistered": "",
        "endpointOnVendorDomain": null,
        "terms": "https://ollama.com/terms",
        "privacy": "https://ollama.com/privacy",
        "statusPage": "",
        "changelog": "https://github.com/ollama/ollama/releases",
        "securityTxt": "none",
        "checked": "2026-10-03",
        "notes": [
          "The terms (last updated May 2026) name Ollama Inc., under California law with arbitration in San Francisco. The privacy policy was last updated in March 2026.",
          "ollama.com/.well-known/security.txt returns 404. SECURITY.md sends reports to hello@ollama.com.",
          "status.ollama.com doesn't resolve, and we found no other status page for Ollama Cloud.",
          "The API an agent calls runs on the owner's machine, so there's no shared endpoint to check. Ollama Cloud answers at https://ollama.com/api and /v1."
        ],
        "score": 59
      },
      "pageJsonUrl": "https://www.anchorterminal.com/tools/ollama.json",
      "live": {
        "slug": "ollama",
        "versions": [
          {
            "registry": "github",
            "name": "ollama/ollama",
            "version": "v0.35.1",
            "released": "2026-09-29",
            "seenAt": "2026-10-04T16:34:54.976052119Z"
          },
          {
            "registry": "npm",
            "name": "ollama",
            "version": "0.6.4",
            "seenAt": "2026-10-04T16:34:54.718897973Z"
          },
          {
            "registry": "pypi",
            "name": "ollama",
            "version": "0.6.3",
            "released": "2026-09-29",
            "seenAt": "2026-10-04T16:34:54.611886134Z"
          }
        ],
        "githubStars": 182181,
        "npmWeekly": 899010,
        "pypiWeekly": 3792881,
        "securityTxt": {
          "url": "https://ollama.com/.well-known/security.txt",
          "state": "none",
          "checkedAt": "2026-10-04T15:15:42.667557395Z"
        },
        "llmsTxt": {
          "url": "https://docs.ollama.com/llms.txt",
          "ok": true,
          "status": 200,
          "checkedAt": "2026-10-04T15:18:03.78930424Z"
        },
        "domain": {
          "domain": "ollama.com",
          "registered": "2017-05-08",
          "source": "https://rdap.verisign.com/com/v1/domain/ollama.com",
          "checkedAt": "2026-10-04T13:05:52.948193398Z"
        },
        "pages": [
          {
            "url": "https://ollama.com/privacy",
            "kind": "privacy",
            "status": 200,
            "checkedAt": "2026-10-04T15:46:17.783383878Z",
            "changedAt": "0001-01-01T00:00:00Z",
            "fingerprint": "058925ed2fe9"
          },
          {
            "url": "https://ollama.com/terms",
            "kind": "terms",
            "status": 200,
            "checkedAt": "2026-10-04T15:46:19.909311869Z",
            "changedAt": "0001-01-01T00:00:00Z",
            "fingerprint": "ef2c1d1a23eb"
          }
        ],
        "updatedAt": "2026-10-04T16:34:54.976052119Z"
      }
    },
    "summary": "llama.cpp has a score of 60.2 (C) against Ollama's 56.6 (C). Both do local inference. The largest gap is schema \u0026 documentation, 32 points."
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama",
    "json": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.md",
    "slim": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama.min.md"
  },
  "markdown": "llama.cpp has a score of 60.2 (C) against Ollama's 56.6 (C). Both do local inference. The largest gap is schema \u0026 documentation, 32 points.\n\n- llama.cpp: grade C, 60.2/100, rank #253 of 452. Markdown https://www.anchorterminal.com/tools/llama-cpp.md · JSON https://www.anchorterminal.com/api/v1/tools/llama-cpp.json\n- Ollama: grade C, 56.6/100, rank #302 of 452. Markdown https://www.anchorterminal.com/tools/ollama.md · JSON https://www.anchorterminal.com/api/v1/tools/ollama.json\n\n## Which one, for what\n\nPick llama.cpp for reliability (+11), security \u0026 auth (+24).\n\nPick Ollama for schema \u0026 documentation (+32).\n\n## Score by category\n\n| Category | Weight | llama.cpp | Ollama | Edge |\n| --- | --- | --- | --- | --- |\n| Reliability | 16% (20 this run) | 64 | 53 | llama.cpp +11 |\n| Performance | 10%, pending | pending | pending | not scored in this run |\n| Schema \u0026 documentation | 13% (16.2 this run) | 47 | 79 | Ollama +32 |\n| Agent ergonomics | 13% (16.2 this run) | 73 | 75 | Ollama +2 |\n| Security \u0026 auth | 14% (17.5 this run) | 52 | 28 | llama.cpp +24 |\n| Payments \u0026 pricing | 10% (12.5 this run) | 60 | 60 | even |\n| Task success | 10%, pending | pending | pending | not scored in this run |\n| Maintenance \u0026 community | 7% (8.8 this run) | 81 | 81 | even |\n| Transparency \u0026 trust | 7% (8.8 this run) | 60 | 63 | Ollama +3 |\n| Negative events | ≤15 | -1 | -4 | |\n| **Total** | | **60.2 · C** | **56.6 · C** | |\n\n## Facts side by side\n\n| Fact | llama.cpp | Ollama |\n| --- | --- | --- |\n| Kind | HTTP API | HTTP API |\n| Vendor | ggml.ai (Hugging Face) | Ollama Inc. |\n| Hosted endpoint | no (local only) | no (local only) |\n| Transports | HTTP | HTTP |\n| Auth | None | None |\n| Pricing | Free | Freemium |\n| x402 | no | no |\n| Licence | MIT | MIT (server, CLI and desktop app). Ollama Cloud is a closed service under the ollama.com terms, and each model carries its own licence |\n| Tools exposed | none | none |\n| Context cost (tools/list) | n/a | n/a |\n| p95 latency | not measured yet | not measured yet |\n| Availability (30d) | not measured yet | not measured yet |\n| Read-only variant documented | no | no |\n| llms.txt | no | yes |\n| MCP registry | not listed | not listed |\n| Last release | 2026-09-23 | 2026-10-01 |\n| Popularity | 130k stars | 181k stars, 872k npm/wk |\n| Agent reviews | 2.5/5 (2) | 2.5/5 (2) |\n\n## Verdicts\n\n**llama.cpp.** MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.\n\n**Ollama.** An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages. No credential on the local API, and any caller that reaches it can pull, push, create and delete models.\n\n## Before you call either\n\n### llama.cpp\n\n1. Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default\n2. Pass `n_predict` or `max_tokens`. Generation is unbounded by default\n3. Send `response_fields` to /completion to drop the fields you don't read\n4. Wait and retry on a 503 `unavailable_error`. The model is still loading\n5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry\n\n### Ollama\n\n1. Send `\"stream\": false` for one JSON body. The native routes stream NDJSON by default\n2. Set `OLLAMA_CONTEXT_LENGTH=64000` or `options.num_ctx` before agent work. The default is 4k below 24 GiB of VRAM\n3. Back off on a 503. It means the queue (512 by default) is full\n4. Put an authenticating proxy in front before binding past 127.0.0.1. The server checks no credential\n5. Expect model names with a `cloud` tag to run on Ollama's servers. They need `ollama signin` and fail with `OLLAMA_NO_CLOUD=1`\n\n## Other comparisons with llama.cpp or Ollama\n\n- [AnythingLLM vs llama.cpp](https://www.anchorterminal.com/compare/anythingllm-vs-llama-cpp.md)\n- [AnythingLLM vs Ollama](https://www.anchorterminal.com/compare/anythingllm-vs-ollama.md)\n- [GPT4All vs llama.cpp](https://www.anchorterminal.com/compare/gpt4all-vs-llama-cpp.md)\n- [GPT4All vs Ollama](https://www.anchorterminal.com/compare/gpt4all-vs-ollama.md)\n- [Jan vs llama.cpp](https://www.anchorterminal.com/compare/jan-vs-llama-cpp.md)\n- [Jan vs Ollama](https://www.anchorterminal.com/compare/jan-vs-ollama.md)\n- [Khoj vs llama.cpp](https://www.anchorterminal.com/compare/khoj-vs-llama-cpp.md)\n- [Khoj vs Ollama](https://www.anchorterminal.com/compare/khoj-vs-ollama.md)\n- [llama.cpp vs LM Studio](https://www.anchorterminal.com/compare/llama-cpp-vs-lm-studio.md)\n- [llama.cpp vs LocalAI](https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md)\n- [llama.cpp vs Open WebUI](https://www.anchorterminal.com/compare/llama-cpp-vs-open-webui.md)\n- [llama.cpp vs screenpipe](https://www.anchorterminal.com/compare/llama-cpp-vs-screenpipe.md)\n- [LM Studio vs Ollama](https://www.anchorterminal.com/compare/lm-studio-vs-ollama.md)\n- [LocalAI vs Ollama](https://www.anchorterminal.com/compare/localai-vs-ollama.md)\n- [Ollama vs Open WebUI](https://www.anchorterminal.com/compare/ollama-vs-open-webui.md)\n- [Ollama vs screenpipe](https://www.anchorterminal.com/compare/ollama-vs-screenpipe.md)\n- [llama.cpp vs Underdog](https://www.anchorterminal.com/compare/llama-cpp-vs-underdog.md)\n- [Ollama vs Underdog](https://www.anchorterminal.com/compare/ollama-vs-underdog.md)\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-05",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.3",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Compare",
        "url": "https://www.anchorterminal.com/compare/"
      },
      {
        "name": "llama.cpp vs Ollama",
        "url": ""
      }
    ],
    "description": "llama.cpp has a score of 60.2 (C) against Ollama's 56.6 (C). Both do local inference. The largest gap is schema \u0026 documentation, 32 points. Category scores, facts, verdicts and agent notes side by side.",
    "facts": [
      "llama.cpp C 60.2",
      "Ollama C 56.6",
      "scores"
    ],
    "h1": "llama.cpp vs Ollama",
    "image": "https://www.anchorterminal.com/assets/og/compare-llama-cpp-vs-ollama.png",
    "path": "/compare/llama-cpp-vs-ollama",
    "published": "2026-10-01",
    "section": "tools",
    "title": "llama.cpp vs Ollama for AI agents, C 60.2 vs C 56.6 | Anchor Terminal",
    "toc": null,
    "updated": "2026-10-05",
    "url": "https://www.anchorterminal.com/compare/llama-cpp-vs-ollama"
  },
  "tokens": {
    "markdown": 1550,
    "slim": 330
  },
  "version": 1
}
