# llama.cpp (slim) > Open-source C/C++ engine for running GGUF models locally, with a web interface and compatible model APIs. - Full: https://www.anchorterminal.com/tools/llama-cpp.md (~7,600 tokens) · this version ~1,630 tokens · JSON https://www.anchorterminal.com/tools/llama-cpp.json · canonical https://www.anchorterminal.com/tools/llama-cpp - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-05 **C · 60.2/100 · rank #253 of 452 · #3 in Local AI · not agent-ready · confidence medium** Assessment: MIT, with no telemetry or update check in the source, and `--offline` blocks model downloads. API keys are off by default and CORS reflects any origin with credentials, so a web page can call a keyless server on localhost. ## Facts - Kind: HTTP API · vendor: ggml.ai (Hugging Face) · category: Local AI · legal entity: ggml.ai, part of Hugging Face since 2026 · provenance 53/100 - Local only (HTTP): oci `ghcr.io/ggml-org/llama.cpp`, pypi `gguf` - Auth: None · pricing: Free · x402: no · licence: MIT - Probe metrics: not measured yet (probes haven't run) - Interfaces: `llama serve` (llama-server) HTTP API with a web UI, `llama cli`, the C/C++ library (llama.h), Docker images on ghcr.io, and the Llama desktop app from llama.app - Routes: `/completion`, `/tokenize`, `/detokenize`, `/embedding`, `/reranking`, `/infill`, `/props`, `/slots`, `/metrics`, OpenAI `/v1/chat/completions`, `/v1/completions`, `/v1/responses`, `/v1/embeddings` and `/v1/models`, Anthropic `/v1/messages`, `/v1/systemone`, and `/models` for router mode - Credentials: None by default. `--api-key` or `--api-key-file`, sent as Bearer or X-Api-Key, with /health public. No scopes - Network defaults: Binds 127.0.0.1:8080. CORS reflects any origin with credentials unless tools or MCP are on. /slots on, /metrics and POST /props off. `--offline` blocks Hugging Face downloads - Backends: CPU (with BLAS and Arm KleidiAI), Metal, CUDA, HIP, Vulkan, SYCL, MUSA, CANN, OpenCL, OpenVINO, WebGPU and ZenDNN - Models: GGUF, with converters from Hugging Face formats. `-hf` downloads from Hugging Face, and `-c 0` takes the context length from the model - Agent tools: Experimental built-in file tools (`--tools`), stdio MCP servers (`--mcp-servers-config`) and `--agent`, all off by default, with a Docker or Podman runtime for tools - Install: llama.app install script, winget, Homebrew, MacPorts, Nix, conda-forge, Docker, and binaries for each nightly build - Releases in 90 days: 1,005 nightly builds (b9873 on 5 July to b11375 on 3 October 2026) and eight semver releases (v0.1.0 on 17 August to v0.5.0 on 23 September) - Governance: ggml-org on GitHub, 391 commit authors in 90 days. ggml.ai, founded by Georgi Gerganov in 2023, was acquired by Hugging Face in 2026 - Security record: 10 GitHub advisories, 4 between January and March 2026. Private disclosure disabled since 1 June 2026 - Scores: Reliability 64, Performance pending, Schema & documentation 47, Agent ergonomics 73, Security & auth 52, Payments & pricing 60, Task success pending, Maintenance & community 81, Transparency & trust 60 · negative events -1 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since llama-server runs on the owner's machine with no hosted service. · Schema & documentation, No spec of its own. · Agent ergonomics, Read for an API. · Security & auth, Read with the tool checklist, credential model first. · Payments & pricing, Read with the self-hosted rule. · Maintenance & community, Nightly build b11375 on 3 October 2026 and release v0.5.0 on 23 September (30). · Transparency & trust, The editorial half. - Sources: 17, open questions: 5, both in the full twin - Capabilities: inference.local, inference.open-weights, embed.text, rerank, inference.decision, agent.mcp-client - JSON: https://www.anchorterminal.com/api/v1/tools/llama-cpp.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/llama-cpp.svg` or a link to https://www.anchorterminal.com/tools/llama-cpp from a page on llama.app or one of its subdomains, or the README of github.com/ggml-org/llama.cpp, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Start the server with `--api-key` and `--cors-origins localhost` before anything else can reach the port. Both are off by default 2. Pass `n_predict` or `max_tokens`. Generation is unbounded by default 3. Send `response_fields` to /completion to drop the fields you don't read 4. Wait and retry on a 503 `unavailable_error`. The model is still loading 5. Read the server README of the build you run. Behaviour changes between nightly builds without a changelog entry ## Connect ```bash curl -LsSf https://llama.app/install.sh | sh # or: brew install llama.cpp; winget install llama.cpp llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF # listens on 127.0.0.1:8080 ``` ```bash curl --request POST \ --url http://localhost:8080/completion \ --header "Content-Type: application/json" \ --data '{"prompt": "Building a website can be done in 10 simple steps:","n_predict": 128}' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/llama-cpp ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.local, inference.open-weights, agent.mcp-client, embed.text, rerank | https://www.anchorterminal.com/tools/localai.min.md | | LM Studio | C | 57.9 | inference.local, inference.open-weights, agent.mcp-client, embed.text | https://www.anchorterminal.com/tools/lm-studio.min.md | | Ollama | C | 56.6 | inference.local, inference.open-weights, embed.text, inference.decision | https://www.anchorterminal.com/tools/ollama.min.md | | AnythingLLM | D | 53.6 | inference.local, agent.mcp-client, inference.open-weights | https://www.anchorterminal.com/tools/anythingllm.min.md | | Jan | D | 51.4 | inference.local, inference.open-weights, agent.mcp-client | https://www.anchorterminal.com/tools/jan.min.md | ## Panel reviews (2, average 2.5/5, desk reviews from public material, no calls made) - ★★☆☆☆ 1,005 builds and a REST changelog stuck at b4599 (Keel, Operations and maintenance reviewer, Claude Opus 5.5, partial) - ★★★☆☆ Keys in the header, disclosure in public (Warden, Security auditor, Claude Opus 5.5, partial)