# Best local AI models and assistants (slim) > LocalAI (B), Lemonade (B) and screenpipe (C) lead the 20 ranked local AI models and assistants. Picks by need, strengths, weaknesses and prices from the Anchor benchmark. - Full: https://www.anchorterminal.com/best/local-ai/index.md (~6,450 tokens) · this version ~1,580 tokens · JSON https://www.anchorterminal.com/best/local-ai/index.json · canonical https://www.anchorterminal.com/best/local-ai/ - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 The 10 highest-scoring of 20 local AI models and assistants on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change. - Ranked: 20 · agent-ready (BB or better): 0 · accept x402: 0 · hosted endpoints: 0 - Full ranked table: https://www.anchorterminal.com/categories/local-ai.md - Head-to-head comparisons: https://www.anchorterminal.com/compare/local-ai/index.md (184) - Methodology: https://www.anchorterminal.com/benchmark/index.md ## The shortlist | # | Tool | Grade | Score | Best for | Price | Where | | --- | --- | --- | --- | --- | --- | --- | | 1 | [LocalAI](https://www.anchorterminal.com/tools/localai.md) | B | 68 | An owner who wants one local server for chat, embeddings, reranking, speech, images and video behind APIs their existing OpenAI, Anthropic or Ollama clients already speak, on almost any accelerator. | Free · OSS | local | | 2 | [Lemonade](https://www.anchorterminal.com/tools/lemonade.md) | B | 63.8 | An owner with AMD hardware (Ryzen AI NPUs, Radeon or Strix Halo) who wants one local server for chat, speech and images that existing OpenAI, Anthropic or Ollama clients can call, and for MCP clients that want local models as tools. | Free · OSS | local | | 3 | [screenpipe](https://www.anchorterminal.com/tools/screenpipe.md) | C | 60.8 | One person who wants an agent to recall what they saw, said or heard on their own computer, meetings included, with local storage and local transcription. | $21 / mo | local | | 4 | [Foundry Local](https://www.anchorterminal.com/tools/foundry-local.md) | C | 60.5 | An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows. | Free · OSS | local | | 5 | [KoboldCpp](https://www.anchorterminal.com/tools/koboldcpp.md) | C | 60.5 | An owner who wants text, image, speech and music models behind one executable with a writing and roleplay interface, and clients that speak the KoboldAI, OpenAI, Ollama or Anthropic formats. | Free · OSS | local | | 6 | [llama.cpp](https://www.anchorterminal.com/tools/llama-cpp.md) | C | 60.2 | An owner who wants the engine itself, any GGUF model, the widest hardware support and the most control over flags, behind an OpenAI- or Anthropic-compatible API. | Free · OSS | local | | 7 | [LM Studio](https://www.anchorterminal.com/tools/lm-studio.md) | C | 57.8 | A machine that serves open models to several agents and tools at once, in whichever API shape each client already speaks, and for headless serving on Linux with llmster. | Free | local | | 8 | [vLLM](https://www.anchorterminal.com/tools/vllm.md) | C | 57.7 | An owner with a GPU server who wants many concurrent requests against one open-weight model behind OpenAI or Anthropic compatible routes. | Free · OSS | local | | 9 | [Docker Model Runner](https://www.anchorterminal.com/tools/docker-model-runner.md) | C | 57.1 | A team that already runs Docker and wants local models served to containers and Compose services through OpenAI-, Anthropic- or Ollama-compatible routes, with models stored as OCI artefacts. | Free · OSS | local | | 10 | [Ollama](https://www.anchorterminal.com/tools/ollama.md) | C | 56.3 | A person or an agent that wants an open model behind a local API with one install, and for pointing Claude Code, Codex or OpenCode at local or cloud models. | $20 / mo | local | ## Picks by need - Highest score overall: [LocalAI](https://www.anchorterminal.com/tools/localai.md), B, 68/100 on the benchmark. Also [Lemonade](https://www.anchorterminal.com/tools/lemonade.md), B, 63.8/100. - Reliability: [Docker Model Runner](https://www.anchorterminal.com/tools/docker-model-runner.md), 85/100 on reliability, against 84 for the overall leader. - Agent ergonomics: [screenpipe](https://www.anchorterminal.com/tools/screenpipe.md), 75/100 on agent ergonomics, against 71 for the overall leader. - Security & auth: [Open WebUI](https://www.anchorterminal.com/tools/open-webui.md), 63/100 on security & auth, against 62 for the overall leader. - Maintenance & community: [Open WebUI](https://www.anchorterminal.com/tools/open-webui.md), 91/100 on maintenance & community, against 80 for the overall leader. - Transparency & trust: [Docker Model Runner](https://www.anchorterminal.com/tools/docker-model-runner.md), 73/100 on transparency & trust, against 47 for the overall leader. Our own products are ranked like any listing but are never a pick. ## How to choose - Network traffic while running: Check what a network monitor sees during model runs and indexing, since outbound traffic during a run may carry prompts or file contents off the machine. - Hardware and memory needs: Check the memory and tokens per second a model reaches on the owner's hardware, since a model that runs slowly will not serve an agent well. - Citing the right source file: Check whether answers cite the file they drew on, since an assistant that cites the wrong document gives an agent confident but unverifiable answers. - Licence for runner and weights: Check the licence on both the runner and the model weights, since an agent deployed for paying clients may fall under terms that restrict use. - How the benchmark tests this category: In this run, public evidence against the published checklist, read as software the owner runs on their own hardware (the local-software lines for Reliability and the self-hosted rule for Payments). When the task suites run, the same small open model, prompts and documents on the same machine through each listing's local API or MCP server. We check the setup steps, tokens per second, memory use, whether answers cite the right file and what a network monitor sees leave the machine. Each listing's verdict, strengths and weaknesses: https://www.anchorterminal.com/best/local-ai/index.md