# KoboldCpp (slim) > Open-source program for running GGUF models on the owner's own computer, built on llama.cpp. One executable serves a web interface and KoboldAI, OpenAI, Ollama and Anthropic compatible APIs on port 5001. - Full: https://www.anchorterminal.com/tools/koboldcpp.md (~7,600 tokens) · this version ~1,930 tokens · JSON https://www.anchorterminal.com/tools/koboldcpp.json · canonical https://www.anchorterminal.com/tools/koboldcpp - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 60.5/100 · rank #462 of 842 · #5 in Local AI · not agent-ready · confidence medium** Assessment: One file runs text, image, speech and music models behind a published OpenAPI 3.0.3 document, with eight releases in 90 days. The server listens on every interface with no password by default, and `--password` leaves the image routes open. ## Facts - Kind: HTTP API · vendor: LostRuins (Concedo) · category: Local AI · legal entity: not named · provenance 27/100 - Local only (HTTP): oci `koboldai/koboldcpp` - Auth: None · pricing: Free · x402: no · licence: AGPL-3.0 for KoboldCpp and KoboldAI Lite. The bundled GGML, llama.cpp and stable-diffusion.cpp code stays under MIT - Probe metrics: not measured yet (probes haven't run) - Interfaces: HTTP server on port 5001 with the KoboldAI Lite web UI at `/`, the llama.cpp web UI at `/lcpp`, StableUI at `/sdui`, a music UI at `/musicui`, interactive API docs at `/api`, a desktop launcher, `--cli` terminal chat and `--agent` - Routes: KoboldAI `/api/v1/generate` and `/api/extra/*` (stream, tokencount, abort, embeddings, transcribe, tts, music, websearch), OpenAI `/v1/chat/completions`, `/v1/completions`, `/v1/responses`, `/v1/embeddings`, `/v1/images/generations`, `/v1/audio/speech`, `/v1/audio/transcriptions` and `/v1/models`, Anthropic `/v1/messages`, Ollama `/api/chat` and `/api/generate`, AUTOMATIC1111 `/sdapi/v1/*`, ComfyUI `/prompt`, and `/mcp` - API contract: OpenAPI 3.0.3 at `/api?json=1` and https://lite.koboldai.net/koboldcpp_api.json, 53 paths and 54 operations. The Ollama, ComfyUI, XTTS and `/mcp` routes are not in it - Credentials: None by default. `--password` or `KCPP_PASSWORD` sets one shared key sent as a Bearer token, for text routes only. `--adminpassword` or `KCPP_ADMINPASSWORD` guards admin routes. No scopes - Network defaults: Listens on all routable interfaces, port 5001. CORS reflects any Origin with credentials. `--ssl` takes a certificate and key. `--remotetunnel` opens a Cloudflare tunnel when asked - Limits: Up to 10 queued requests by default (`--multiuser`), `--ratelimit` seconds between requests per IP (off by default), `--maxrequestsize` 32 MB, `--genlimit` caps output. Busy and rate-limited requests get 503 - Models: GGUF text models, legacy GGML `.bin`, vision projectors, Stable Diffusion, SDXL, SD3, Flux and other image models, Whisper, several TTS models, ACE Step music models and embedding models, per the README. No model is bundled - Backends: CPU, CUDA, Vulkan and Metal in the release binaries. ROCm in a rolling Linux build - Install: Single-file binaries on GitHub releases for Windows x64, Linux x64 and Apple Silicon macOS, each with `nocuda` and `oldpc` variants where relevant, the `koboldai/koboldcpp` Docker image, a Colab notebook, and source builds for Android (Termux), OpenBSD and Raspberry Pi - Agent and tools: `--agent` starts a terminal agent with nine tools and confirmation modes on, auto and off. `--mcpfile` connects stdio and HTTP MCP servers. `--websearch` turns on a DuckDuckGo proxy. All off by default - Releases in 90 days: Eight, from 1.117.1 on 10 July to 1.122.1 on 27 September 2026, each with notes on GitHub - Governance: Maintained by LostRuins (Concedo), the main author of the commits on the `concedo` branch that are not upstream llama.cpp work. The Windows binary names KoboldAI as the company. No legal entity is stated - Security record: One GitHub advisory, GHSA-qhvp-gj7g-rw26 (Low, 29 August 2026). No `SECURITY.md`. Private vulnerability reporting is open on GitHub - Scores: Reliability 68, Performance pending, Schema & documentation 68, Agent ergonomics 63, Security & auth 38, Payments & pricing 60, Task success pending, Maintenance & community 82, Transparency & trust 49 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since KoboldCpp runs on the owner's machine with no hosted service. · Schema & documentation, An OpenAPI 3.0.3 document is served by the program at `/api?json=1` and published at lite.koboldai.net, with 53 paths and 54 operations. · Agent ergonomics, Read for an API. · Security & auth, Read with the tool checklist, credential model first. · Payments & pricing, Read with the self-hosted rule. · Maintenance & community, Release 1.122.1 on 27 September 2026, 11 days before the check (30). · Transparency & trust, The editorial half. - Sources: 18, open questions: 8, both in the full twin - Capabilities: inference.local, inference.open-weights, agent.mcp-client, embed.text, image.generate, speech.stt, speech.tts, audio.music - JSON: https://www.anchorterminal.com/api/v1/tools/koboldcpp.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/koboldcpp.svg` or a link to https://www.anchorterminal.com/tools/koboldcpp from a page on koboldcpp.net or one of its subdomains, or the README of github.com/LostRuins/koboldcpp, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Start with `--host 127.0.0.1` and `--password`. The default listens on every interface with no key 2. Send the password as `Authorization: Bearer `. It is not read from the query string 3. Treat 503 as both busy and rate limited. The server never sends 429 or `Retry-After`, and the wait in seconds is in `detail.msg` 4. Pass `max_length` or `max_tokens`. The default is 2,048 tokens unless `--defaultgenamt` changes it 5. Send a `genkey` with each generation so `/api/extra/generate/check` and `/api/extra/abort` act on your request and not another caller's ## Connect ```bash curl -fLo koboldcpp-linux-x64 https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp-linux-x64 && chmod +x koboldcpp-linux-x64 && ./koboldcpp-linux-x64 ./koboldcpp-linux-x64 --model /path/to/model.gguf # listens on port 5001 ``` ```bash curl --request POST \ --url http://localhost:5001/api/v1/generate \ --header "Content-Type: application/json" \ --data '{"prompt": "Niko the kobold stalked carefully down the alley,", "max_context_length": 2048, "max_length": 100}' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/koboldcpp ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.local, inference.open-weights, agent.mcp-client, embed.text, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/localai.min.md | | Lemonade | B | 63.8 | inference.local, inference.open-weights, embed.text, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/lemonade.min.md | | DeepInfra | B | 63 | inference.open-weights, embed.text, image.generate, speech.stt, speech.tts | https://www.anchorterminal.com/tools/deepinfra.min.md | | TextGen | E | 45.1 | inference.local, inference.open-weights, agent.mcp-client, embed.text, image.generate | https://www.anchorterminal.com/tools/text-generation-webui.min.md | | Foundry Local | C | 60.5 | inference.local, inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/foundry-local.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)