{
  "data": {
    "category": {
      "area": "models",
      "capabilities": [
        "inference.local",
        "inference.open-weights",
        "memory.user",
        "memory.search",
        "agent.mcp-client"
      ],
      "description": "Software that runs models on hardware the owner keeps, a laptop, a desktop or a home server. Model runners with a local API, chat apps, and personal assistants that work from the owner's own files, mail and records, with no cloud account needed. Compared on what models they run, what hardware they need, what leaves the machine, what an agent can call and the licence.",
      "json": "https://www.anchorterminal.com/categories/local-ai.json",
      "name": "Local AI",
      "slug": "local-ai",
      "test": "In this run, public evidence against the published checklist, read as software the owner runs on their own hardware (the local-software lines for Reliability and the self-hosted rule for Payments). When the task suites run, the same small open model, prompts and documents on the same machine through each listing's local API or MCP server. We check the setup steps, tokens per second, memory use, whether answers cite the right file and what a network monitor sees leave the machine.",
      "title": "Local AI: models and assistants that run on your own hardware",
      "toolCount": 20,
      "tools": [
        "localai",
        "lemonade",
        "screenpipe",
        "foundry-local",
        "koboldcpp",
        "llama-cpp",
        "lm-studio",
        "vllm",
        "docker-model-runner",
        "ollama",
        "anythingllm",
        "localghost",
        "mlx-lm",
        "open-webui",
        "jan",
        "text-generation-webui",
        "khoj",
        "gpt4all",
        "underdog",
        "ghost-core"
      ],
      "url": "https://www.anchorterminal.com/categories/local-ai"
    },
    "faq": [
      {
        "answer": "LocalAI has the highest benchmark score of the 20 ranked local AI models and assistants, 68 (B). Lemonade is second with 63.8 (B).",
        "question": "What are the highest-rated local AI models and assistants for AI agents?"
      },
      {
        "answer": "0 of the 20 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.",
        "question": "How many local AI models and assistants are agent-ready?"
      },
      {
        "answer": "None of the ranked listings here accepts x402 for its main call yet.",
        "question": "Which local AI models and assistants accept x402 payments?"
      },
      {
        "answer": "By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.",
        "question": "How is this list ranked?"
      }
    ],
    "howToChoose": [
      {
        "label": "Network traffic while running",
        "detail": "Check what a network monitor sees during model runs and indexing, since outbound traffic during a run may carry prompts or file contents off the machine."
      },
      {
        "label": "Hardware and memory needs",
        "detail": "Check the memory and tokens per second a model reaches on the owner's hardware, since a model that runs slowly will not serve an agent well."
      },
      {
        "label": "Citing the right source file",
        "detail": "Check whether answers cite the file they drew on, since an assistant that cites the wrong document gives an agent confident but unverifiable answers."
      },
      {
        "label": "Licence for runner and weights",
        "detail": "Check the licence on both the runner and the model weights, since an agent deployed for paying clients may fall under terms that restrict use."
      }
    ],
    "picks": [
      {
        "also": {
          "name": "Lemonade",
          "slug": "lemonade",
          "why": "B, 63.8/100"
        },
        "name": "LocalAI",
        "need": "Highest score overall",
        "slug": "localai",
        "why": "B, 68/100 on the benchmark"
      },
      {
        "name": "Docker Model Runner",
        "need": "Reliability",
        "slug": "docker-model-runner",
        "why": "85/100 on reliability, against 84 for the overall leader"
      },
      {
        "name": "screenpipe",
        "need": "Agent ergonomics",
        "slug": "screenpipe",
        "why": "75/100 on agent ergonomics, against 71 for the overall leader"
      },
      {
        "name": "Open WebUI",
        "need": "Security \u0026 auth",
        "slug": "open-webui",
        "why": "63/100 on security \u0026 auth, against 62 for the overall leader"
      },
      {
        "name": "Open WebUI",
        "need": "Maintenance \u0026 community",
        "slug": "open-webui",
        "why": "91/100 on maintenance \u0026 community, against 80 for the overall leader"
      },
      {
        "name": "Docker Model Runner",
        "need": "Transparency \u0026 trust",
        "slug": "docker-model-runner",
        "why": "73/100 on transparency \u0026 trust, against 47 for the overall leader"
      }
    ],
    "ranked": 20,
    "shortlist": [
      {
        "bestFor": "An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator.",
        "grade": "B",
        "name": "LocalAI",
        "position": 1,
        "price": "Free · OSS",
        "score": 68,
        "slug": "localai",
        "strengths": [
          "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app",
          "OpenAI, Anthropic, Open Responses, Ollama and ElevenLabs-compatible endpoints, with a Swagger 2.0 file of 133 operations served by every instance",
          "429 and 503 responses carry Retry-After, and errors come in the calling client's own envelope"
        ],
        "url": "https://www.anchorterminal.com/tools/localai",
        "verdict": "MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.",
        "weaknesses": [
          "No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin",
          "CVE-2026-59707, an unauthenticated SSRF in v4.3.1 and earlier, published by VulnCheck in July 2026 with no advisory from the project",
          "SECURITY.md still names 3.x as the supported series, and there's no security.txt or privacy policy"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "An owner with AMD hardware (Ryzen AI NPUs, Radeon or Strix Halo) who wants one local server for chat, speech and images that existing OpenAI, Anthropic or Ollama clients can call, and for MCP clients that want local models as tools.",
        "grade": "B",
        "name": "Lemonade",
        "position": 2,
        "price": "Free · OSS",
        "score": 63.8,
        "slug": "lemonade",
        "strengths": [
          "Apache 2.0, with OpenAI, Anthropic Messages, Ollama and llama.cpp-compatible routes on one port (13305)",
          "`POST /mcp` exposes six tools over Streamable HTTP, and omitting `model` reuses a loaded or downloaded model before any download",
          "Every running server returns its own Markdown API reference at `GET /v1/docs` and through the `lemonade_docs` tool, so it matches the installed version"
        ],
        "url": "https://www.anchorterminal.com/tools/lemonade",
        "verdict": "Apache 2.0, with weekly releases, installers for Windows, macOS and five Linux routes, and a six-tool MCP endpoint whose descriptions steer a caller away from accidental multi-gigabyte downloads. Authentication is off by default, the repository has no security policy, and the docs tell WebSocket clients to pass the key in the URL.",
        "weaknesses": [
          "No authentication by default. With no key set, every route answers, including `/internal/*` shutdown and configuration",
          "GitHub reports no `SECURITY.md`, and we found no security.txt, advisory or disclosure address",
          "The docs tell WebSocket clients to send the key as `?api_key=KEY` in the URL"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "One person who wants an agent to recall what they saw, said or heard on their own computer, meetings included, with local storage and local transcription.",
        "grade": "C",
        "name": "screenpipe",
        "position": 3,
        "price": "$21 / mo",
        "score": 60.8,
        "slug": "screenpipe",
        "strengths": [
          "33 MCP tools with typed JSON Schemas, every one annotated, 21 marked read-only and `merge-speakers` marked destructive",
          "`limit` and `offset`, time, app, window, speaker and tag filters, and per-result truncation at 1,000 characters by default",
          "The local API needs a key on every request by default, localhost included, and pipes get their own permission-limited tokens"
        ],
        "url": "https://www.anchorterminal.com/tools/screenpipe",
        "verdict": "33 MCP tools with typed JSON Schemas, every one annotated, 21 marked read-only and `merge-speakers` marked destructive. 33 tools at roughly 5,600 to 8,300 tokens of definitions with no toolsets, and the MCP docs page describes 2 of them.",
        "weaknesses": [
          "33 tools at roughly 5,600 to 8,300 tokens of definitions with no toolsets, and the MCP docs page describes 2 of them",
          "The local key is `sp-` plus 8 hexadecimal characters, and the docs list passing it as a `?token=` query parameter",
          "PostHog analytics, Sentry and, since 17 September 2026, remote support logs are on by default, and the README said nothing was sent to external servers until 1 October"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.",
        "grade": "C",
        "name": "Foundry Local",
        "position": 4,
        "price": "Free · OSS",
        "score": 60.5,
        "slug": "foundry-local",
        "strengths": [
          "SDK 2.1.0 published on npm, PyPI, NuGet and crates.io on 29 September 2026, with the SDK and the v2 native runtime under MIT in a public repository",
          "A model alias selects the best variant for the machine's CPU, GPU or NPU, with CUDA, WebGPU, OpenVINO, QNN and Vitis AI execution providers",
          "The v2 server answers `/v1/chat/completions`, `/v1/responses`, `/v1/embeddings`, `/v1/audio/transcriptions` and `/v1/models` in OpenAI's shapes"
        ],
        "url": "https://www.anchorterminal.com/tools/foundry-local",
        "verdict": "The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.",
        "weaknesses": [
          "The local server takes no credential, and its routes include model load and unload and `POST /shutdown`",
          "The REST reference on Microsoft Learn lists `/openai/*` and `/foundry/list` routes that the v2 runtime source doesn't register, and doesn't cover `/v1/responses`",
          "Telemetry is on by default through Microsoft's 1DS SDK. The opt-out covers non-essential telemetry only, and the Learn FAQ doesn't mention it"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "An owner who wants text, image, speech and music models behind one executable with a writing and roleplay interface, and clients that speak the KoboldAI, OpenAI, Ollama or Anthropic formats.",
        "grade": "C",
        "name": "KoboldCpp",
        "position": 5,
        "price": "Free · OSS",
        "score": 60.5,
        "slug": "koboldcpp",
        "strengths": [
          "OpenAPI 3.0.3 document with 54 operations, served by the program at `/api?json=1` and published at lite.koboldai.net",
          "KoboldAI, OpenAI, Ollama, Anthropic, AUTOMATIC1111 and ComfyUI style routes from one server on port 5001",
          "Eight releases between 10 July and 27 September 2026, each with written notes, and replies on all 24 of the newest open issues"
        ],
        "url": "https://www.anchorterminal.com/tools/koboldcpp",
        "verdict": "One file runs text, image, speech and music models behind a published OpenAPI 3.0.3 document, with eight releases in 90 days. The server listens on every interface with no password by default, and `--password` leaves the image routes open.",
        "weaknesses": [
          "With no `--host` the server accepts connections on all routable interfaces, and no password is set by default",
          "`--password` covers text routes only. The `--help` text says image endpoints are not secured",
          "CORS reflects any Origin with credentials allowed and permits private-network requests"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.",
        "grade": "C",
        "name": "llama.cpp",
        "position": 6,
        "price": "Free · OSS",
        "score": 60.2,
        "slug": "llama-cpp",
        "strengths": [
          "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads",
          "OpenAI chat completions, responses and embeddings, Anthropic messages, reranking and /v1/systemone from one server",
          "`response_fields`, `json_schema` and `grammar` control the size and shape of output, and errors carry an OpenAI-style type and code"
        ],
        "url": "https://www.anchorterminal.com/tools/llama-cpp",
        "verdict": "MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.",
        "weaknesses": [
          "API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost",
          "No OpenAPI file of its own, and the REST API changelog stops at b4599",
          "Private security disclosure disabled since 1 June 2026, with fixes asked for as public pull requests"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "A machine that serves open models to several agents and tools at once, in whichever API shape each client already speaks, and for headless serving on Linux with llmster.",
        "grade": "C",
        "name": "LM Studio",
        "position": 7,
        "price": "Free",
        "score": 57.8,
        "slug": "lm-studio",
        "strengths": [
          "OpenAI-compatible chat completions, responses, completions and embeddings, Anthropic-compatible /v1/messages and a native /api/v1, all on one port",
          "llmster, a headless daemon installed with one command, with a documented systemd setup for Linux servers",
          "Named API tokens with permissions, and API access to MCP servers behind two switches, one of which also needs authentication on"
        ],
        "url": "https://www.anchorterminal.com/tools/lm-studio",
        "verdict": "OpenAI-compatible chat completions, responses, completions and embeddings, Anthropic-compatible /v1/messages and a native /api/v1, all on one port. Authentication is off by default, so any local process can call the server.",
        "weaknesses": [
          "Authentication is off by default, so any local process can call the server",
          "Closed-source app and daemon with no public CI or test suite",
          "The Python SDK's last stable release (1.5.0, 22 August 2025) can't send API tokens, and the docs name an environment variable no release reads"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.",
        "grade": "C",
        "name": "vLLM",
        "position": 8,
        "price": "Free · OSS",
        "score": 57.7,
        "slug": "vllm",
        "strengths": [
          "OpenAI chat, completions, responses and embeddings, Anthropic `/v1/messages`, Cohere embed and rerank, transcription and `/v1/systemone` from one server",
          "Apache-2.0, with a written three-stage deprecation policy and release notes that carry a breaking changes section",
          "Eight stable releases between 12 July and 2 October 2026, and v0.31.0 lists 717 commits from 307 contributors"
        ],
        "url": "https://www.anchorterminal.com/tools/vllm",
        "verdict": "Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.",
        "weaknesses": [
          "`--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it",
          "No key by default, the server binds every interface when `--host` is unset, and CORS allows any origin",
          "At least 81 GitHub security advisories in 12 months, two critical, most of them remote crashes or resource exhaustion"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "A team that already runs Docker and wants local models served to containers and Compose services through OpenAI-, Anthropic- or Ollama-compatible routes, with models stored as OCI artefacts.",
        "grade": "C",
        "name": "Docker Model Runner",
        "position": 9,
        "price": "Free · OSS",
        "score": 57.1,
        "slug": "docker-model-runner",
        "strengths": [
          "OpenAI-, Anthropic- and Ollama-compatible routes on one local port, so existing clients for those three APIs work with a changed base URL",
          "CI runs lint, race-detector tests and end-to-end tests on every push to main, and the ten most recent runs on main passed on 8 October 2026",
          "Two GitHub security advisories with CVE numbers, fixed versions and workarounds, and a SECURITY.md that promises an acknowledgement within 72 hours"
        ],
        "url": "https://www.anchorterminal.com/tools/docker-model-runner",
        "verdict": "CI passes on the main branch, and Docker has published two security advisories with CVEs and fixed versions for the project. The API takes no credential, so any client or container that reaches it can pull, delete and run models, and the documentation has no OpenAPI file or error reference.",
        "weaknesses": [
          "No credential on the API. The docs say any client that can reach it, including other containers, can pull, load and run models",
          "No OpenAPI file, no error reference and no rate-limit or retry guidance in the reviewed documentation",
          "Two releases in the 90 days to 8 October 2026 (v1.2.7 and v1.2.8), the latest on 12 August"
        ],
        "where": "local",
        "x402": "no"
      },
      {
        "bestFor": "A person or an agent that wants an open model behind a local API with one install, and for pointing Claude Code, Codex or OpenCode at local or cloud models.",
        "grade": "C",
        "name": "Ollama",
        "position": 10,
        "price": "$20 / mo",
        "score": 56.3,
        "slug": "ollama",
        "strengths": [
          "An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages",
          "Native, OpenAI-compatible and Anthropic-compatible routes on one local port, with `ollama launch` for Claude Code, Codex and OpenCode",
          "28 releases in the 90 days to 3 October 2026, and official Python and JavaScript libraries released on 28 September"
        ],
        "url": "https://www.anchorterminal.com/tools/ollama",
        "verdict": "An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages. No credential on the local API, and any caller that reaches it can pull, push, create and delete models.",
        "weaknesses": [
          "No credential on the local API, and any caller that reaches it can pull, push, create and delete models",
          "No GitHub security advisory, against 12 CVEs on NVD since October 2025",
          "The Windows updater installed unsigned files until v0.23.3 on 12 May 2026, fixed under a release note that didn't mention security"
        ],
        "where": "local",
        "x402": "no"
      }
    ],
    "updated": "2026-10-09"
  },
  "kind": "anchor.page",
  "links": {
    "api": "https://www.anchorterminal.com/api/v1/index.json",
    "html": "https://www.anchorterminal.com/best/local-ai/",
    "json": "https://www.anchorterminal.com/best/local-ai/index.json",
    "llms": "https://www.anchorterminal.com/llms.txt",
    "markdown": "https://www.anchorterminal.com/best/local-ai/index.md",
    "slim": "https://www.anchorterminal.com/best/local-ai/index.min.md"
  },
  "markdown": "The 10 highest-scoring of 20 local AI models and assistants on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.\n\n- Ranked: 20 · agent-ready (BB or better): 0 · accept x402: 0 · hosted endpoints: 0\n- Full ranked table: https://www.anchorterminal.com/categories/local-ai.md\n- Head-to-head comparisons: https://www.anchorterminal.com/compare/local-ai/index.md (184)\n- Methodology: https://www.anchorterminal.com/benchmark/index.md\n\n## The shortlist\n\n| # | Tool | Grade | Score | Best for | Price | Where |\n| --- | --- | --- | --- | --- | --- | --- |\n| 1 | [LocalAI](https://www.anchorterminal.com/tools/localai.md) | B | 68 | An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator. | Free · OSS | local |\n| 2 | [Lemonade](https://www.anchorterminal.com/tools/lemonade.md) | B | 63.8 | An owner with AMD hardware (Ryzen AI NPUs, Radeon or Strix Halo) who wants one local server for chat, speech and images that existing OpenAI, Anthropic or Ollama clients can call, and for MCP clients that want local models as tools. | Free · OSS | local |\n| 3 | [screenpipe](https://www.anchorterminal.com/tools/screenpipe.md) | C | 60.8 | One person who wants an agent to recall what they saw, said or heard on their own computer, meetings included, with local storage and local transcription. | $21 / mo | local |\n| 4 | [Foundry Local](https://www.anchorterminal.com/tools/foundry-local.md) | C | 60.5 | An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows. | Free · OSS | local |\n| 5 | [KoboldCpp](https://www.anchorterminal.com/tools/koboldcpp.md) | C | 60.5 | An owner who wants text, image, speech and music models behind one executable with a writing and roleplay interface, and clients that speak the KoboldAI, OpenAI, Ollama or Anthropic formats. | Free · OSS | local |\n| 6 | [llama.cpp](https://www.anchorterminal.com/tools/llama-cpp.md) | C | 60.2 | An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API. | Free · OSS | local |\n| 7 | [LM Studio](https://www.anchorterminal.com/tools/lm-studio.md) | C | 57.8 | A machine that serves open models to several agents and tools at once, in whichever API shape each client already speaks, and for headless serving on Linux with llmster. | Free | local |\n| 8 | [vLLM](https://www.anchorterminal.com/tools/vllm.md) | C | 57.7 | An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes. | Free · OSS | local |\n| 9 | [Docker Model Runner](https://www.anchorterminal.com/tools/docker-model-runner.md) | C | 57.1 | A team that already runs Docker and wants local models served to containers and Compose services through OpenAI-, Anthropic- or Ollama-compatible routes, with models stored as OCI artefacts. | Free · OSS | local |\n| 10 | [Ollama](https://www.anchorterminal.com/tools/ollama.md) | C | 56.3 | A person or an agent that wants an open model behind a local API with one install, and for pointing Claude Code, Codex or OpenCode at local or cloud models. | $20 / mo | local |\n\n## Picks by need\n\n- Highest score overall: [LocalAI](https://www.anchorterminal.com/tools/localai.md), B, 68/100 on the benchmark. Also [Lemonade](https://www.anchorterminal.com/tools/lemonade.md), B, 63.8/100.\n- Reliability: [Docker Model Runner](https://www.anchorterminal.com/tools/docker-model-runner.md), 85/100 on reliability, against 84 for the overall leader.\n- Agent ergonomics: [screenpipe](https://www.anchorterminal.com/tools/screenpipe.md), 75/100 on agent ergonomics, against 71 for the overall leader.\n- Security \u0026 auth: [Open WebUI](https://www.anchorterminal.com/tools/open-webui.md), 63/100 on security \u0026 auth, against 62 for the overall leader.\n- Maintenance \u0026 community: [Open WebUI](https://www.anchorterminal.com/tools/open-webui.md), 91/100 on maintenance \u0026 community, against 80 for the overall leader.\n- Transparency \u0026 trust: [Docker Model Runner](https://www.anchorterminal.com/tools/docker-model-runner.md), 73/100 on transparency \u0026 trust, against 47 for the overall leader.\n\nOur own products are ranked like any listing but are never a pick.\n\n## How to choose\n\n- Network traffic while running: Check what a network monitor sees during model runs and indexing, since outbound traffic during a run may carry prompts or file contents off the machine.\n- Hardware and memory needs: Check the memory and tokens per second a model reaches on the owner's hardware, since a model that runs slowly will not serve an agent well.\n- Citing the right source file: Check whether answers cite the file they drew on, since an assistant that cites the wrong document gives an agent confident but unverifiable answers.\n- Licence for runner and weights: Check the licence on both the runner and the model weights, since an agent deployed for paying clients may fall under terms that restrict use.\n\n- How the benchmark tests this category: In this run, public evidence against the published checklist, read as software the owner runs on their own hardware (the local-software lines for Reliability and the self-hosted rule for Payments). When the task suites run, the same small open model, prompts and documents on the same machine through each listing's local API or MCP server. We check the setup steps, tokens per second, memory use, whether answers cite the right file and what a network monitor sees leave the machine.\n\n## Each one in detail\n\n### 1. LocalAI, B 68/100\n\nOpen-source engine in Go, MIT licensed, that runs models on the owner's hardware behind OpenAI-, Anthropic-, Ollama- and ElevenLabs-compatible APIs on port 8080.\n\n- Verdict: MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin.\n- Choose it for: An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator.\n- Strength: MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app\n- Strength: OpenAI, Anthropic, Open Responses, Ollama and ElevenLabs-compatible endpoints, with a Swagger 2.0 file of 133 operations served by every instance\n- Strength: 429 and 503 responses carry Retry-After, and errors come in the calling client's own envelope\n- Weakness: No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin\n- Weakness: CVE-2026-59707, an unauthenticated SSRF in v4.3.1 and earlier, published by VulnCheck in July 2026 with no advisory from the project\n- Weakness: SECURITY.md still names 3.x as the supported series, and there's no security.txt or privacy policy\n- Price: Free · OSS · Auth: OAuth or key · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/localai.md\n\n### 2. Lemonade, B 63.8/100\n\nOpen-source local AI server from AMD and community contributors. It runs text, speech and image models on the owner's CPU, GPU or NPU behind OpenAI-, Anthropic- and Ollama-compatible APIs and an MCP endpoint on port 13305.\n\n- Verdict: Apache 2.0, with weekly releases, installers for Windows, macOS and five Linux routes, and a six-tool MCP endpoint whose descriptions steer a caller away from accidental multi-gigabyte downloads. Authentication is off by default, the repository has no security policy, and the docs tell WebSocket clients to pass the key in the URL.\n- Choose it for: An owner with AMD hardware (Ryzen AI NPUs, Radeon or Strix Halo) who wants one local server for chat, speech and images that existing OpenAI, Anthropic or Ollama clients can call, and for MCP clients that want local models as tools.\n- Strength: Apache 2.0, with OpenAI, Anthropic Messages, Ollama and llama.cpp-compatible routes on one port (13305)\n- Strength: `POST /mcp` exposes six tools over Streamable HTTP, and omitting `model` reuses a loaded or downloaded model before any download\n- Strength: Every running server returns its own Markdown API reference at `GET /v1/docs` and through the `lemonade_docs` tool, so it matches the installed version\n- Weakness: No authentication by default. With no key set, every route answers, including `/internal/*` shutdown and configuration\n- Weakness: GitHub reports no `SECURITY.md`, and we found no security.txt, advisory or disclosure address\n- Weakness: The docs tell WebSocket clients to send the key as `?api_key=KEY` in the URL\n- Price: Free · OSS · Auth: API key · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/lemonade.md\n- Against #1: https://www.anchorterminal.com/compare/lemonade-vs-localai.md\n\n### 3. screenpipe, C 60.8/100\n\nDesktop app and CLI from Negentropy Labs, Inc. (Screenpipe, YC S26) that records the owner's screen and audio continuously on macOS, Windows and Linux.\n\n- Verdict: 33 MCP tools with typed JSON Schemas, every one annotated, 21 marked read-only and `merge-speakers` marked destructive. 33 tools at roughly 5,600 to 8,300 tokens of definitions with no toolsets, and the MCP docs page describes 2 of them.\n- Choose it for: One person who wants an agent to recall what they saw, said or heard on their own computer, meetings included, with local storage and local transcription.\n- Strength: 33 MCP tools with typed JSON Schemas, every one annotated, 21 marked read-only and `merge-speakers` marked destructive\n- Strength: `limit` and `offset`, time, app, window, speaker and tag filters, and per-result truncation at 1,000 characters by default\n- Strength: The local API needs a key on every request by default, localhost included, and pipes get their own permission-limited tokens\n- Weakness: 33 tools at roughly 5,600 to 8,300 tokens of definitions with no toolsets, and the MCP docs page describes 2 of them\n- Weakness: The local key is `sp-` plus 8 hexadecimal characters, and the docs list passing it as a `?token=` query parameter\n- Weakness: PostHog analytics, Sentry and, since 17 September 2026, remote support logs are on by default, and the README said nothing was sent to external servers until 1 October\n- Price: $21 / mo · Auth: API key · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/screenpipe.md\n- Against #1: https://www.anchorterminal.com/compare/localai-vs-screenpipe.md\n- Disclosure: Screenpipe competes with LocalGhost, which Anchor Terminal's founder builds, and LocalGhost's own about page names it as a competitor. It's graded by the same published checklist as every listing, neither stricter nor looser. Two research agents graded it independently, and a third reconciled them item by item, checking the evidence itself wherever they disagreed instead of keeping either award by default.\n\n### 4. Foundry Local, C 60.5/100\n\nMicrosoft's on-device model runtime, built on ONNX Runtime. Applications embed it through SDKs for C#, JavaScript, Python and Rust, and it can start an optional OpenAI-compatible server on localhost. A preview CLI is also available.\n\n- Verdict: The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.\n- Choose it for: An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.\n- Strength: SDK 2.1.0 published on npm, PyPI, NuGet and crates.io on 29 September 2026, with the SDK and the v2 native runtime under MIT in a public repository\n- Strength: A model alias selects the best variant for the machine's CPU, GPU or NPU, with CUDA, WebGPU, OpenVINO, QNN and Vitis AI execution providers\n- Strength: The v2 server answers `/v1/chat/completions`, `/v1/responses`, `/v1/embeddings`, `/v1/audio/transcriptions` and `/v1/models` in OpenAI's shapes\n- Weakness: The local server takes no credential, and its routes include model load and unload and `POST /shutdown`\n- Weakness: The REST reference on Microsoft Learn lists `/openai/*` and `/foundry/list` routes that the v2 runtime source doesn't register, and doesn't cover `/v1/responses`\n- Weakness: Telemetry is on by default through Microsoft's 1DS SDK. The opt-out covers non-essential telemetry only, and the Learn FAQ doesn't mention it\n- Price: Free · OSS · Auth: None · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/foundry-local.md\n- Against #1: https://www.anchorterminal.com/compare/foundry-local-vs-localai.md\n\n### 5. KoboldCpp, C 60.5/100\n\nOpen-source program for running GGUF models on the owner's own computer, built on llama.cpp. One executable serves a web interface and KoboldAI, OpenAI, Ollama and Anthropic compatible APIs on port 5001.\n\n- Verdict: One file runs text, image, speech and music models behind a published OpenAPI 3.0.3 document, with eight releases in 90 days. The server listens on every interface with no password by default, and `--password` leaves the image routes open.\n- Choose it for: An owner who wants text, image, speech and music models behind one executable with a writing and roleplay interface, and clients that speak the KoboldAI, OpenAI, Ollama or Anthropic formats.\n- Strength: OpenAPI 3.0.3 document with 54 operations, served by the program at `/api?json=1` and published at lite.koboldai.net\n- Strength: KoboldAI, OpenAI, Ollama, Anthropic, AUTOMATIC1111 and ComfyUI style routes from one server on port 5001\n- Strength: Eight releases between 10 July and 27 September 2026, each with written notes, and replies on all 24 of the newest open issues\n- Weakness: With no `--host` the server accepts connections on all routable interfaces, and no password is set by default\n- Weakness: `--password` covers text routes only. The `--help` text says image endpoints are not secured\n- Weakness: CORS reflects any Origin with credentials allowed and permits private-network requests\n- Price: Free · OSS · Auth: None · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/koboldcpp.md\n- Against #1: https://www.anchorterminal.com/compare/koboldcpp-vs-localai.md\n\n### 6. llama.cpp, C 60.2/100\n\nOpen-source C/C++ engine for running GGUF models locally, with a web interface and compatible model APIs.\n\n- Verdict: MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost.\n- Choose it for: An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API.\n- Strength: MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads\n- Strength: OpenAI chat completions, responses and embeddings, Anthropic messages, reranking and /v1/systemone from one server\n- Strength: `response_fields`, `json_schema` and `grammar` control the size and shape of output, and errors carry an OpenAI-style type and code\n- Weakness: API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost\n- Weakness: No OpenAPI file of its own, and the REST API changelog stops at b4599\n- Weakness: Private security disclosure disabled since 1 June 2026, with fixes asked for as public pull requests\n- Price: Free · OSS · Auth: None · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/llama-cpp.md\n- Against #1: https://www.anchorterminal.com/compare/llama-cpp-vs-localai.md\n\n### 7. LM Studio, C 57.8/100\n\nDesktop app and headless daemon from Element Labs for running open-weight models on the owner's machine with llama.cpp and MLX, plus the Splash engine on Apple silicon M3 or newer since 0.4.25.\n\n- Verdict: OpenAI-compatible chat completions, responses, completions and embeddings, Anthropic-compatible /v1/messages and a native /api/v1, all on one port. Authentication is off by default, so any local process can call the server.\n- Choose it for: A machine that serves open models to several agents and tools at once, in whichever API shape each client already speaks, and for headless serving on Linux with llmster.\n- Strength: OpenAI-compatible chat completions, responses, completions and embeddings, Anthropic-compatible /v1/messages and a native /api/v1, all on one port\n- Strength: llmster, a headless daemon installed with one command, with a documented systemd setup for Linux servers\n- Strength: Named API tokens with permissions, and API access to MCP servers behind two switches, one of which also needs authentication on\n- Weakness: Authentication is off by default, so any local process can call the server\n- Weakness: Closed-source app and daemon with no public CI or test suite\n- Weakness: The Python SDK's last stable release (1.5.0, 22 August 2025) can't send API tokens, and the docs name an environment variable no release reads\n- Price: Free · Auth: API key · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/lm-studio.md\n- Against #1: https://www.anchorterminal.com/compare/lm-studio-vs-localai.md\n\n### 8. vLLM, C 57.7/100\n\nvLLM is an open-source inference and serving engine for open-weight language models. `vllm serve` runs an HTTP server with OpenAI-compatible, Anthropic Messages, embedding, reranking and transcription routes on the owner's own GPUs or CPUs.\n\n- Verdict: Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026.\n- Choose it for: An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes.\n- Strength: OpenAI chat, completions, responses and embeddings, Anthropic `/v1/messages`, Cohere embed and rerank, transcription and `/v1/systemone` from one server\n- Strength: Apache-2.0, with a written three-stage deprecation policy and release notes that carry a breaking changes section\n- Strength: Eight stable releases between 12 July and 2 October 2026, and v0.31.0 lists 717 commits from 307 contributors\n- Weakness: `--api-key` guards only the `/v1`, `/v2`, `/inference` and `/cohere` prefixes. `/invocations`, `/pooling`, `/classify`, `/score`, `/rerank`, `/pause` and `/update_weights` answer without it\n- Weakness: No key by default, the server binds every interface when `--host` is unset, and CORS allows any origin\n- Weakness: At least 81 GitHub security advisories in 12 months, two critical, most of them remote crashes or resource exhaustion\n- Price: Free · OSS · Auth: None · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/vllm.md\n- Against #1: https://www.anchorterminal.com/compare/localai-vs-vllm.md\n\n### 9. Docker Model Runner, C 57.1/100\n\nDocker's open-source tool for pulling and running open models from Docker Hub, OCI registries or Hugging Face. It runs through Docker Desktop, Docker Engine or a standalone `dmr` binary, with local OpenAI-, Anthropic- and Ollama-compatible APIs.\n\n- Verdict: CI passes on the main branch, and Docker has published two security advisories with CVEs and fixed versions for the project. The API takes no credential, so any client or container that reaches it can pull, delete and run models, and the documentation has no OpenAPI file or error reference.\n- Choose it for: A team that already runs Docker and wants local models served to containers and Compose services through OpenAI-, Anthropic- or Ollama-compatible routes, with models stored as OCI artefacts.\n- Strength: OpenAI-, Anthropic- and Ollama-compatible routes on one local port, so existing clients for those three APIs work with a changed base URL\n- Strength: CI runs lint, race-detector tests and end-to-end tests on every push to main, and the ten most recent runs on main passed on 8 October 2026\n- Strength: Two GitHub security advisories with CVE numbers, fixed versions and workarounds, and a SECURITY.md that promises an acknowledgement within 72 hours\n- Weakness: No credential on the API. The docs say any client that can reach it, including other containers, can pull, load and run models\n- Weakness: No OpenAPI file, no error reference and no rate-limit or retry guidance in the reviewed documentation\n- Weakness: Two releases in the 90 days to 8 October 2026 (v1.2.7 and v1.2.8), the latest on 12 August\n- Price: Free · OSS · Auth: None · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/docker-model-runner.md\n- Against #1: https://www.anchorterminal.com/compare/docker-model-runner-vs-localai.md\n\n### 10. Ollama, C 56.3/100\n\nOpen-source model runner for macOS, Windows and Linux, with a local API and a library of downloadable models.\n\n- Verdict: An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages. No credential on the local API, and any caller that reaches it can pull, push, create and delete models.\n- Choose it for: A person or an agent that wants an open model behind a local API with one install, and for pointing Claude Code, Codex or OpenCode at local or cloud models.\n- Strength: An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages\n- Strength: Native, OpenAI-compatible and Anthropic-compatible routes on one local port, with `ollama launch` for Claude Code, Codex and OpenCode\n- Strength: 28 releases in the 90 days to 3 October 2026, and official Python and JavaScript libraries released on 28 September\n- Weakness: No credential on the local API, and any caller that reaches it can pull, push, create and delete models\n- Weakness: No GitHub security advisory, against 12 CVEs on NVD since October 2025\n- Weakness: The Windows updater installed unsigned files until v0.23.3 on 12 May 2026, fixed under a release note that didn't mention security\n- Price: $20 / mo · Auth: None · x402: no · Where: local\n- Full assessment: https://www.anchorterminal.com/tools/ollama.md\n- Against #1: https://www.anchorterminal.com/compare/localai-vs-ollama.md\n\n10 more are ranked in the full table: https://www.anchorterminal.com/categories/local-ai.md\n\n## Head to head\n\n- [Lemonade vs LocalAI](https://www.anchorterminal.com/compare/lemonade-vs-localai.md)\n- [LocalAI vs screenpipe](https://www.anchorterminal.com/compare/localai-vs-screenpipe.md)\n- [Foundry Local vs LocalAI](https://www.anchorterminal.com/compare/foundry-local-vs-localai.md)\n- [KoboldCpp vs LocalAI](https://www.anchorterminal.com/compare/koboldcpp-vs-localai.md)\n- [Lemonade vs screenpipe](https://www.anchorterminal.com/compare/lemonade-vs-screenpipe.md)\n- [Foundry Local vs Lemonade](https://www.anchorterminal.com/compare/foundry-local-vs-lemonade.md)\n- [KoboldCpp vs Lemonade](https://www.anchorterminal.com/compare/koboldcpp-vs-lemonade.md)\n- [Foundry Local vs screenpipe](https://www.anchorterminal.com/compare/foundry-local-vs-screenpipe.md)\n- [KoboldCpp vs screenpipe](https://www.anchorterminal.com/compare/koboldcpp-vs-screenpipe.md)\n- [Foundry Local vs KoboldCpp](https://www.anchorterminal.com/compare/foundry-local-vs-koboldcpp.md)\n\n## Questions\n\n### What are the highest-rated local AI models and assistants for AI agents?\n\nLocalAI has the highest benchmark score of the 20 ranked local AI models and assistants, 68 (B). Lemonade is second with 63.8 (B).\n\n### How many local AI models and assistants are agent-ready?\n\n0 of the 20 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.\n\n### Which local AI models and assistants accept x402 payments?\n\nNone of the ranked listings here accepts x402 for its main call yet.\n\n### How is this list ranked?\n\nBy the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 9 October 2026.\n\n## How this list is made\n\nThe order is the Anchor benchmark score, the same number as on each listing. Each listing is graded from public evidence against the benchmark checklist, and the picks are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.\n",
  "meta": {
    "attribution": "Anchor Terminal (https://www.anchorterminal.com)",
    "docs": "https://www.anchorterminal.com/docs/",
    "generatedAt": "2026-10-10",
    "license": "CC-BY-4.0",
    "method": "https://www.anchorterminal.com/benchmark/",
    "methodology": "0.4",
    "openapi": "https://www.anchorterminal.com/openapi.json",
    "preview": false,
    "run": "2026-10-01",
    "runLabel": "October 2026 research run"
  },
  "page": {
    "breadcrumbs": [
      {
        "name": "Home",
        "url": "https://www.anchorterminal.com/"
      },
      {
        "name": "Best of",
        "url": "https://www.anchorterminal.com/best/"
      },
      {
        "name": "Local AI",
        "url": ""
      }
    ],
    "description": "LocalAI (B), Lemonade (B) and screenpipe (C) lead the 20 ranked local AI models and assistants. Picks by need, strengths, weaknesses and prices from the Anchor benchmark.",
    "facts": [
      "LocalAI B",
      "Lemonade B",
      "screenpipe C"
    ],
    "h1": "Best local AI models and assistants",
    "image": "https://www.anchorterminal.com/assets/og/best-local-ai.png",
    "path": "/best/local-ai/",
    "published": "",
    "section": "tools",
    "title": "Best local AI models and assistants in 2026, ranked | Anchor Terminal",
    "toc": null,
    "updated": "2026-10-09",
    "url": "https://www.anchorterminal.com/best/local-ai/"
  },
  "tokens": {
    "markdown": 6450,
    "slim": 1580
  },
  "version": 1
}
