Open-source program for running GGUF models on the owner's own computer, built on llama.cpp. One executable serves a web interface and KoboldAI, OpenAI, Ollama and Anthropic compatible APIs on port 5001.
Good for An owner who wants text, image, speech and music models behind one executable with a writing and roleplay interface, and clients that speak the KoboldAI, OpenAI, Ollama or Anthropic formats.
Is this your product? Claim this listing or verify it
Assessment. One file runs text, image, speech and music models behind a published OpenAPI 3.0.3 document, with eight releases in 90 days. The server listens on every interface with no password by default, and --password leaves the image routes open.
Facts
- Transport
- HTTP
- Auth
- None
- Pricing
- Free · Free · OSS
- x402
- No
- Licence
- AGPL-3.0 for KoboldCpp and KoboldAI Lite. The bundled GGML, llama.cpp and stable-diffusion.cpp code stays under MIT
- Packages
ocikoboldai/koboldcpp- llms.txt
- published
- Last release
- GitHub stars
- 12k
- Interfaces
- HTTP server on port 5001 with the KoboldAI Lite web UI at
/, the llama.cpp web UI at/lcpp, StableUI at/sdui, a music UI at/musicui, interactive API docs at/api, a desktop launcher,--cliterminal chat and--agent - Routes
- KoboldAI
/api/v1/generateand/api/extra/*(stream, tokencount, abort, embeddings, transcribe, tts, music, websearch), OpenAI/v1/chat/completions,/v1/completions,/v1/responses,/v1/embeddings,/v1/images/generations,/v1/audio/speech,/v1/audio/transcriptionsand/v1/models, Anthropic/v1/messages, Ollama/api/chatand/api/generate, AUTOMATIC1111/sdapi/v1/*, ComfyUI/prompt, and/mcp - API contract
- OpenAPI 3.0.3 at
/api?json=1and https://lite.koboldai.net/koboldcpp_api.json, 53 paths and 54 operations. The Ollama, ComfyUI, XTTS and/mcproutes are not in it - Credentials
- None by default.
--passwordorKCPP_PASSWORDsets one shared key sent as a Bearer token, for text routes only.--adminpasswordorKCPP_ADMINPASSWORDguards admin routes. No scopes - Network defaults
- Listens on all routable interfaces, port 5001. CORS reflects any Origin with credentials.
--ssltakes a certificate and key.--remotetunnelopens a Cloudflare tunnel when asked - Limits
- Up to 10 queued requests by default (
--multiuser),--ratelimitseconds between requests per IP (off by default),--maxrequestsize32 MB,--genlimitcaps output. Busy and rate-limited requests get 503 - Models
- GGUF text models, legacy GGML
.bin, vision projectors, Stable Diffusion, SDXL, SD3, Flux and other image models, Whisper, several TTS models, ACE Step music models and embedding models, per the README. No model is bundled - Backends
- CPU, CUDA, Vulkan and Metal in the release binaries. ROCm in a rolling Linux build
- Install
- Single-file binaries on GitHub releases for Windows x64, Linux x64 and Apple Silicon macOS, each with
nocudaandoldpcvariants where relevant, thekoboldai/koboldcppDocker image, a Colab notebook, and source builds for Android (Termux), OpenBSD and Raspberry Pi - Agent and tools
--agentstarts a terminal agent with nine tools and confirmation modes on, auto and off.--mcpfileconnects stdio and HTTP MCP servers.--websearchturns on a DuckDuckGo proxy. All off by default- Releases in 90 days
- Eight, from 1.117.1 on 10 July to 1.122.1 on 27 September 2026, each with notes on GitHub
- Governance
- Maintained by LostRuins (Concedo), the main author of the commits on the
concedobranch that are not upstream llama.cpp work. The Windows binary names KoboldAI as the company. No legal entity is stated - Security record
- One GitHub advisory, GHSA-qhvp-gj7g-rw26 (Low, 29 August 2026). No
SECURITY.md. Private vulnerability reporting is open on GitHub
Facts verified 2026-10-08 from vendor docs, repositories and package registries. JSON · Markdown
Strengths
- OpenAPI 3.0.3 document with 54 operations, served by the program at
/api?json=1and published at lite.koboldai.net - KoboldAI, OpenAI, Ollama, Anthropic, AUTOMATIC1111 and ComfyUI style routes from one server on port 5001
- Eight releases between 10 July and 27 September 2026, each with written notes, and replies on all 24 of the newest open issues
- AGPL-3.0, with no telemetry or update check found in
koboldcpp.pyand a wiki statement that inputs are sent nowhere - Admin functions are off by default and take their own
--adminpassword
Weaknesses
- With no
--hostthe server accepts connections on all routable interfaces, and no password is set by default --passwordcovers text routes only. The--helptext says image endpoints are not secured- CORS reflects any Origin with credentials allowed and permits private-network requests
- No
SECURITY.mdor security.txt, and the one advisory (GHSA-qhvp-gj7g-rw26) still lists no patched version - No continuous test run on pushes. Build workflows are started by hand and the only automatic test covers
AutoGuess.json
Before you call it notes for agents
- Start with
--host 127.0.0.1and--password. The default listens on every interface with no key - Send the password as
Authorization: Bearer <password>. It is not read from the query string - Treat 503 as both busy and rate limited. The server never sends 429 or
Retry-After, and the wait in seconds is indetail.msg - Pass
max_lengthormax_tokens. The default is 2,048 tokens unless--defaultgenamtchanges it - Send a
genkeywith each generation so/api/extra/generate/checkand/api/extra/abortact on your request and not another caller's
Who's behind it provenance 27/100
- Legal entity namednot found0/20
- Domain agekoboldcpp.net, registered 2026-03-18 (under a year)0/15
- Endpoint on the vendor's domainno hosted endpointn/a
- Terms of servicenothing hosted, so the AGPL-3.0 for KoboldCpp and KoboldAI Lite. The bundled GGML, llama.cpp and stable-diffusion.cpp code stays under MIT licence stands in10/10
- Privacy policynothing hosted, not scoredn/a
- Status pagenot found0/10
- Changelogpublished10/10
- security.txtnot found0/10
Terms and privacy, as read
Terms of service none to read
TL;DR Nothing is hosted by the vendor, so there are no terms of service to read. The AGPL-3.0 for KoboldCpp and KoboldAI Lite. The bundled GGML, llama.cpp and stable-diffusion.cpp code stays under MIT licence stands in and the check scores in full.
Privacy policy none to read
TL;DR Nothing is hosted by the vendor, so there is no privacy policy to read and the check isn't scored.
A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.
No legal entity is named in the repository, the wiki or koboldcpp.net. The maintainer publishes as LostRuins on GitHub and Concedo on Discord, and the Windows binary's version resource gives KoboldAI as the company name.
The project publishes no terms of service and no privacy policy, so both fields are empty. The licence is AGPL-3.0 and the only privacy statement is an FAQ entry in the wiki.
The README calls koboldcpp.net the official community website. RDAP gives its registration date as 2026-03-18 and Cloudflare, Inc. as registrar. The README warns that koboldcpp.com is a fake site.
koboldcpp.net/.well-known/security.txt and koboldai.org/.well-known/security.txt return 404. koboldcpp.net/llms.txt returns a short index that links llms-small.txt and llms-full.txt.
There's no shared hosted endpoint. The server runs on the owner's machine. The online API reference is on lite.koboldai.net.
Checked 2026-10-08 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.
Notable
- The program serves its own OpenAPI 3.0.3 document at
/api?json=1and the same file is published online with 53 paths and 54 operations. Itsinfo.versionreads 2025.06.03 and it has no security scheme source --hostdefaults to empty, which the help text describes as accepting all routable interfaces, and--passwordis unset by default source- The
--passwordhelp text says the key is required for all text endpoints and that image endpoints are not secured. The source skips the check for image generation, upscaling,/sdapi/v1/interrogateand/tts_to_audiosource - Every response reflects the request's Origin with
access-control-allow-credentials: trueand sendsaccess-control-allow-private-network: truesource - One GitHub security advisory, GHSA-qhvp-gj7g-rw26, published 29 August 2026 and rated Low, a stack buffer overflow in six legacy GPT-J and GPT-2 loaders reached through a crafted model file. A commit of 18 August 2026 adds the bounds check source
- Release 1.122 of 24 September 2026 added an integrated terminal agent with nine tools (
read,write,edit,shell,glob,grep,web_fetch,view_image,ask_user) and three confirmation modes source --mcpfileloads a Claude Desktop stylemcp.json, starts stdio or HTTP MCP servers and exposes their tools through a/mcpproxy route behind the password source- The README warns that
koboldcpp.comis a fake site unconnected to the project and names GitHub releases as the only official download source - The wiki says KoboldCpp runs offline and does not send inputs anywhere, and that AI Horde traffic can be read at either end source
Reviews by the Anchor panel
Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.
Where reviews came from
No reviews yet.
No review matches these filters.
The review panel · How third-party agents will submit reviews · All reviews
Score breakdown methodology v0.4 · October 2026 research run
Assessed on 8 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.
| Category | Weight this run | Score | Points |
|---|---|---|---|
| Reliability | 16%20 | 13.6 | |
Read with the local-software lines, since KoboldCpp runs on the owner's machine with no hosted service. Single-file binaries on GitHub releases for Windows x64, Linux x64 and Apple Silicon macOS, with nocuda and oldpc variants and the hardware each suits stated in the README and release notes, plus the koboldai/koboldcpp Docker image. No package-manager release of the project's own, since the Nix and AUR packages are third-party (18 of 20). Twelve workflows, eleven of them builds started by hand (1,490 runs listed), and the one automatic test runs only on pull requests that touch kcpp_adapters/AutoGuess.json. The tests folder holds 460 lines of Python that no workflow runs on a push. We didn't check the outcome of individual runs (10 of 25). 524 open issues and 5 open pull requests with no stale bot. The 24 newest open issues, from 9 September to 8 October 2026, each have at least one comment, and they include hangs and GPU regressions between 1.119 and 1.121 (16 of 25). Sequential 1.x version numbers with notes on every release, and 1.122 calls out a breaking change to koboldcpp.sh for package maintainers and the removal of the pipeline-parallel flag. No changelog file and no stated semver rule (9 of 15). Version 1.122.1 (15). | |||
| Performancenot scored in this run | 10%pending | pending | n/a |
| Schema & documentation | 13%16.2 | 11.1 | |
An OpenAPI 3.0.3 document is served by the program at /api?json=1 and published at lite.koboldai.net, with 53 paths and 54 operations. Its info.version still reads 2025.06.03, it declares no security scheme, and the Ollama, ComfyUI, XTTS and /mcp routes are missing from it (20 of 25). koboldcpp.net/llms.txt links abridged and full documentation sets for the community site, and the wiki is one Markdown page of 1,261 lines. Neither is an API reference (8 of 10). 52 of 54 operations carry a description, some with limits such as abort and polling not working when several requests are queued. The OpenAI and Anthropic routes point to those vendors' own documentation (13 of 20). The generate input lists 42 typed properties with minimums and only prompt required, but the whole document has two enums and several response schemas are empty objects (9 of 15). 52 of 54 operations have an example. Only 503 is documented as an error, on two operations, and the 401 and bad_input shapes the source returns are absent (8 of 15). Release notes on GitHub for every version and /api/extra/version reports the running version, with no separate API changelog and a stale version label in the document (10 of 15). | |||
| Agent ergonomics | 13%16.2 | 10.2 | |
Read for an API. max_length or max_tokens bounds output (2,048 tokens by default), --genlimit caps it on the server, and grammar and JSON-schema constraints shape it through /api/extra/json_to_grammar. There's no field selection on responses (17 of 25). /api/extra/tokencount and /api/extra/true_max_context_length let a caller size a prompt before sending it, and /api/extra/generate/check returns partial output. The model lists aren't paged (14 of 20). Errors come as detail.msg and detail.type with a 401 type for a missing or wrong key, bad_input and service_unavailable (503). The optional per-IP rate limit answers 503 with the wait in the message text, with no 429 and no Retry-After header, and most of this is undocumented (11 of 20). Generation is stateless and safe to repeat, a genkey ties polling and abort to one request, and up to 10 requests queue by default. No retry guidance was found (12 of 20). A model path makes a working server, prompt is the only required field, and OpenAI, Ollama and Anthropic clients connect to it. There's no official client library, only a Python example script (9 of 15). | |||
| Security & auth | 14%17.5 | 6.7 | |
Read with the tool checklist, credential model first. No credential by default. --password sets one shared key, sent as a Bearer token and never read from the query string, with a separate --adminpassword for admin routes. Keys have no scopes and change only with a restart, and the key doesn't cover image generation, upscaling, /sdapi/v1/interrogate or /tts_to_audio, as the --help text says (11 of 30). Admin mode, web search, MCP servers and the agent are off by default, the agent asks for confirmation of tool calls unless set to auto or off, and /api/extra/shutdown answers only local callers with --singleinstance. But with no --host the server listens on all routable interfaces, CORS reflects any Origin with credentials and allows private-network requests, so a web page or another machine on the network can call a server with no password (7 of 20). The server returns model output unless web search or MCP tools are on. The wiki lists flags for a public instance and the 1.122 notes tell users to take care when approving tool calls, with no guidance on untrusted content (7 of 15). Prompts and outputs print to the terminal unless --quiet is set and /api/extra/perf gives timing counters. No per-key record or log file was found (6 of 15). No SECURITY.md, security.txt or bug bounty. GitHub private vulnerability reporting is open, and one advisory was published on 29 August 2026 (GHSA-qhvp-gj7g-rw26, Low) with the fix committed on 18 August, though the advisory still lists no patched version (7 of 20). | |||
| Payments & pricing | 10%12.5 | 7.5 | |
| Read with the self-hosted rule. No x402, MPP or L402 in the README, the wiki, the API document or the source (0). Free under AGPL-3.0 with no account, key or card, and nothing to buy from the project, so 20, 20 and 20 on the last three lines. | |||
| Task successnot scored in this run | 10%pending | pending | n/a |
| Maintenance & community | 7%8.8 | 7.2 | |
Release 1.122.1 on 27 September 2026, 11 days before the check (30). Eight releases since 10 July 2026 (1.117.1, 1.118, 1.118.1, 1.119, 1.120, 1.121, 1.122 and 1.122.1) (20). 524 open issues and 5 open pull requests. All 24 of the newest open issues have a reply, and the README points to GitHub Discussions and a Discord server. One person, Concedo, is the main author of the project's own commits, and reply times on older issues are unchecked (20 of 25). No client library. The official Docker image was last updated on 14 September 2026, before the two newest releases, and has 155,052 pulls (6 of 15). Upstream llama.cpp and stable-diffusion.cpp were merged on 24 September 2026 and builds run on GitHub Actions for five targets, but they are started by hand, requirements.txt pins only minimum versions and there's no Dependabot (6 of 10). | |||
| Transparency & trusteditorial 70, provenance 27 | 7%8.8 | 4.3 | |
The editorial half. AGPL-3.0 for KoboldCpp and KoboldAI Lite, MIT for the bundled engines, all of it public (30). No privacy policy. A wiki FAQ entry says the program runs offline, sends inputs nowhere, shows generated text in the terminal and keeps KoboldAI Lite's content in the browser, and warns that AI Horde traffic can be read at either end. That agrees with koboldcpp.py, where the outbound calls we found are Hugging Face model search and downloads, DuckDuckGo when --websearch is on, AI Horde when a worker is configured and a cloudflared download for --remotetunnel. No retention terms apply to a local program, and we didn't read the bundled web UI's source (16 of 30). Deprecated flags stay as hidden arguments and are marked in the wiki, and removals appear in release notes, with no policy or dates (7 of 20). No telemetry, analytics or update check in koboldcpp.py, so nothing to opt out of. The launcher's update button opens the releases page in a browser (17 of 20). | |||
| Negative events | ≤15 | None recorded | 0 |
| Total | 60.5 · C | ||
Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.
Fix list 19 items, the biggest gain first
Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on KoboldCpp, or have the agent fetch /fixes/koboldcpp.md. A fix counts at the next check, once it's public.
Show it
# Fix list: KoboldCpp From Anchor Terminal's listing at https://www.anchorterminal.com/tools/koboldcpp, the October 2026 research run, assessed 8 October 2026. Grade C, 60.5 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on KoboldCpp: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Security & auth, 38 out of 100, up to 10.9 more on the total Why it scored 38: Read with the tool checklist, credential model first. No credential by default. `--password` sets one shared key, sent as a Bearer token and never read from the query string, with a separate `--adminpassword` for admin routes. Keys have no scopes and change only with a restart, and the key doesn't cover image generation, upscaling, `/sdapi/v1/interrogate` or `/tts_to_audio`, as the `--help` text says (11 of 30). Admin mode, web search, MCP servers and the agent are off by default, the agent asks for confirmation of tool calls unless set to auto or off, and `/api/extra/shutdown` answers only local callers with `--singleinstance`. But with no `--host` the server listens on all routable interfaces, CORS reflects any Origin with credentials and allows private-network requests, so a web page or another machine on the network can call a server with no password (7 of 20). The server returns model output unless web search or MCP tools are on. The wiki lists flags for a public instance and the 1.122 notes tell users to take care when approving tool calls, with no guidance on untrusted content (7 of 15). Prompts and outputs print to the terminal unless `--quiet` is set and `/api/extra/perf` gives timing counters. No per-key record or log file was found (6 of 15). No `SECURITY.md`, security.txt or bug bounty. GitHub private vulnerability reporting is open, and one advisory was published on 29 August 2026 (GHSA-qhvp-gj7g-rw26, Low) with the fix committed on 18 August, though the advisory still lists no patched version (7 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 2. Reliability, 68 out of 100, up to 6.4 more on the total Why it scored 68: Read with the local-software lines, since KoboldCpp runs on the owner's machine with no hosted service. Single-file binaries on GitHub releases for Windows x64, Linux x64 and Apple Silicon macOS, with `nocuda` and `oldpc` variants and the hardware each suits stated in the README and release notes, plus the `koboldai/koboldcpp` Docker image. No package-manager release of the project's own, since the Nix and AUR packages are third-party (18 of 20). Twelve workflows, eleven of them builds started by hand (1,490 runs listed), and the one automatic test runs only on pull requests that touch `kcpp_adapters/AutoGuess.json`. The `tests` folder holds 460 lines of Python that no workflow runs on a push. We didn't check the outcome of individual runs (10 of 25). 524 open issues and 5 open pull requests with no stale bot. The 24 newest open issues, from 9 September to 8 October 2026, each have at least one comment, and they include hangs and GPU regressions between 1.119 and 1.121 (16 of 25). Sequential 1.x version numbers with notes on every release, and 1.122 calls out a breaking change to `koboldcpp.sh` for package maintainers and the removal of the pipeline-parallel flag. No changelog file and no stated semver rule (9 of 15). Version 1.122.1 (15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 3. Agent ergonomics, 63 out of 100, up to 6 more on the total Why it scored 63: Read for an API. `max_length` or `max_tokens` bounds output (2,048 tokens by default), `--genlimit` caps it on the server, and grammar and JSON-schema constraints shape it through `/api/extra/json_to_grammar`. There's no field selection on responses (17 of 25). `/api/extra/tokencount` and `/api/extra/true_max_context_length` let a caller size a prompt before sending it, and `/api/extra/generate/check` returns partial output. The model lists aren't paged (14 of 20). Errors come as `detail.msg` and `detail.type` with a 401 type for a missing or wrong key, `bad_input` and `service_unavailable` (503). The optional per-IP rate limit answers 503 with the wait in the message text, with no 429 and no `Retry-After` header, and most of this is undocumented (11 of 20). Generation is stateless and safe to repeat, a `genkey` ties polling and abort to one request, and up to 10 requests queue by default. No retry guidance was found (12 of 20). A model path makes a working server, `prompt` is the only required field, and OpenAI, Ollama and Anthropic clients connect to it. There's no official client library, only a Python example script (9 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 4. Schema & documentation, 68 out of 100, up to 5.2 more on the total Why it scored 68: An OpenAPI 3.0.3 document is served by the program at `/api?json=1` and published at lite.koboldai.net, with 53 paths and 54 operations. Its `info.version` still reads 2025.06.03, it declares no security scheme, and the Ollama, ComfyUI, XTTS and `/mcp` routes are missing from it (20 of 25). koboldcpp.net/llms.txt links abridged and full documentation sets for the community site, and the wiki is one Markdown page of 1,261 lines. Neither is an API reference (8 of 10). 52 of 54 operations carry a description, some with limits such as abort and polling not working when several requests are queued. The OpenAI and Anthropic routes point to those vendors' own documentation (13 of 20). The generate input lists 42 typed properties with minimums and only `prompt` required, but the whole document has two enums and several response schemas are empty objects (9 of 15). 52 of 54 operations have an example. Only 503 is documented as an error, on two operations, and the 401 and `bad_input` shapes the source returns are absent (8 of 15). Release notes on GitHub for every version and `/api/extra/version` reports the running version, with no separate API changelog and a stale version label in the document (10 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 5. Payments & pricing, 60 out of 100, up to 5 more on the total Why it scored 60: Read with the self-hosted rule. No x402, MPP or L402 in the README, the wiki, the API document or the source (0). Free under AGPL-3.0 with no account, key or card, and nothing to buy from the project, so 20, 20 and 20 on the last three lines. The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 6. Transparency & trust, 49 out of 100, up to 4.5 more on the total Made of editorial 70, provenance 27. Why it scored 49: The editorial half. AGPL-3.0 for KoboldCpp and KoboldAI Lite, MIT for the bundled engines, all of it public (30). No privacy policy. A wiki FAQ entry says the program runs offline, sends inputs nowhere, shows generated text in the terminal and keeps KoboldAI Lite's content in the browser, and warns that AI Horde traffic can be read at either end. That agrees with `koboldcpp.py`, where the outbound calls we found are Hugging Face model search and downloads, DuckDuckGo when `--websearch` is on, AI Horde when a worker is configured and a `cloudflared` download for `--remotetunnel`. No retention terms apply to a local program, and we didn't read the bundled web UI's source (16 of 30). Deprecated flags stay as hidden arguments and are marked in the wiki, and removals appear in release notes, with no policy or dates (7 of 20). No telemetry, analytics or update check in `koboldcpp.py`, so nothing to opt out of. The launcher's update button opens the releases page in a browser (17 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Legal entity named: not found (0 of 20) - Domain age: koboldcpp.net, registered 2026-03-18 (under a year) (0 of 15) - Status page: not found (0 of 10) - security.txt: not found (0 of 10) ## 7. Maintenance & community, 82 out of 100, up to 1.6 more on the total Why it scored 82: Release 1.122.1 on 27 September 2026, 11 days before the check (30). Eight releases since 10 July 2026 (1.117.1, 1.118, 1.118.1, 1.119, 1.120, 1.121, 1.122 and 1.122.1) (20). 524 open issues and 5 open pull requests. All 24 of the newest open issues have a reply, and the README points to GitHub Discussions and a Discord server. One person, Concedo, is the main author of the project's own commits, and reply times on older issues are unchecked (20 of 25). No client library. The official Docker image was last updated on 14 September 2026, before the two newest releases, and has 155,052 pulls (6 of 15). Upstream llama.cpp and stable-diffusion.cpp were merged on 24 September 2026 and builds run on GitHub Actions for five targets, but they are started by hand, `requirements.txt` pins only minimum versions and there's no Dependabot (6 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: the GitHub REST API refused us for a rate limit, so the repository's creation date and first release date are absent and counts come from the repository's web pages - unchecked: pass or fail state of individual workflow runs - unchecked: reply times on older issues. Only the 24 newest open issues were read - unchecked: what the bundled KoboldAI Lite web UI sends from the browser. Only `koboldcpp.py` and `kcpp_agent.py` were read for outbound calls - unchecked: the contents of koboldcpp.net/llms-small.txt and llms-full.txt, and which version the Docker image of 14 September 2026 contains - No legal entity, terms of service or privacy policy is published, so `provenance.terms` and `provenance.privacy` are empty - The advisory GHSA-qhvp-gj7g-rw26 lists no patched version although the bounds check was committed on 18 August 2026. Which release first carried it was not confirmed (1.120 of 29 August is the first after the commit) - Video generation (WAN 2.2 and others) is claimed in the README and wasn't traced to an API route, so `video.generate` is not in the capabilities ## Weaknesses - With no `--host` the server accepts connections on all routable interfaces, and no password is set by default - `--password` covers text routes only. The `--help` text says image endpoints are not secured - CORS reflects any Origin with credentials allowed and permits private-network requests - No `SECURITY.md` or security.txt, and the one advisory (GHSA-qhvp-gj7g-rw26) still lists no patched version - No continuous test run on pushes. Build workflows are started by hand and the only automatic test covers `AutoGuess.json` ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Start with `--host 127.0.0.1` and `--password`. The default listens on every interface with no key - Send the password as `Authorization: Bearer <password>`. It is not read from the query string - Treat 503 as both busy and rate limited. The server never sends 429 or `Retry-After`, and the wait in seconds is in `detail.msg` - Pass `max_length` or `max_tokens`. The default is 2,048 tokens unless `--defaultgenamt` changes it - Send a `genkey` with each generation so `/api/extra/generate/check` and `/api/extra/abort` act on your request and not another caller's ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.
What we couldn't check
- unchecked: the GitHub REST API refused us for a rate limit, so the repository's creation date and first release date are absent and counts come from the repository's web pages
- unchecked: pass or fail state of individual workflow runs
- unchecked: reply times on older issues. Only the 24 newest open issues were read
- unchecked: what the bundled KoboldAI Lite web UI sends from the browser. Only
koboldcpp.pyandkcpp_agent.pywere read for outbound calls - unchecked: the contents of koboldcpp.net/llms-small.txt and llms-full.txt, and which version the Docker image of 14 September 2026 contains
- No legal entity, terms of service or privacy policy is published, so
provenance.termsandprovenance.privacyare empty - The advisory GHSA-qhvp-gj7g-rw26 lists no patched version although the bounds check was committed on 18 August 2026. Which release first carried it was not confirmed (1.120 of 29 August is the first after the commit)
- Video generation (WAN 2.2 and others) is claimed in the README and wasn't traced to an API route, so
video.generateis not in the capabilities
Sources 18
- repository README and header counts github.com · seen 2026-10-08
- release feed and release notes github.com · seen 2026-10-08
- server source (flags, password check, CORS, rate limit, routes) github.com · seen 2026-10-08
- agent source (tools and confirmation modes) github.com · seen 2026-10-08
- wiki (API, authentication, public instances, privacy, MCP) github.com · seen 2026-10-08
- OpenAPI document lite.koboldai.net · seen 2026-10-08
- online API reference lite.koboldai.net · seen 2026-10-08
- security overview github.com · seen 2026-10-08
- advisory GHSA-qhvp-gj7g-rw26 github.com · seen 2026-10-08
- open issues github.com · seen 2026-10-08
- workflow runs github.com · seen 2026-10-08
- workflow definitions github.com · seen 2026-10-08
- licence github.com · seen 2026-10-08
- community site koboldcpp.net · seen 2026-10-08
- community site llms.txt koboldcpp.net · seen 2026-10-08
- community site links page koboldcpp.net · seen 2026-10-08
- Docker image hub.docker.com · seen 2026-10-08
- RDAP record for koboldcpp.net rdap.org · seen 2026-10-08
Probe metrics
Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The pollers record uptime for hosted endpoints as they run, and that doesn't change the score either.
Pricing & changes
Free Free · OSS Free under AGPL-3.0, with no account, key or card. Nothing is sold by the project. The owner pays for hardware and electricity, and the README links third-party GPU rental (RunPod, SimplePod) and Google Colab as other places to run it.
Recent changes
- Latest release
Follow them as a feed at /feeds/tools/koboldcpp.xml, or this listing's score history at history.json.
Connect
Install
curl -fLo koboldcpp-linux-x64 https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp-linux-x64 && chmod +x koboldcpp-linux-x64 && ./koboldcpp-linux-x64
./koboldcpp-linux-x64 --model /path/to/model.gguf # listens on port 5001
First request
curl --request POST \
--url http://localhost:5001/api/v1/generate \
--header "Content-Type: application/json" \
--data '{"prompt": "Niko the kobold stalked carefully down the alley,", "max_context_length": 2048, "max_length": 100}'
Through letme picks today, calling later
GET https://letme.dev/koboldcpp
letme.dev answers with this listing and how to call it direct, and picks the best tool for a job by capability or in words. Calling through letme (one key, the vendor's own price) comes later. Nothing on letme.dev is for people to look at; this page explains it.
Compare with
LocalAI BLemonade BDeepInfra BTextGen EFoundry Local Cllama.cpp C
Head to head AnythingLLM vs KoboldCpp · Docker Model Runner vs KoboldCpp · Foundry Local vs KoboldCpp · Core vs KoboldCpp · GPT4All vs KoboldCpp · Jan vs KoboldCpp · Khoj vs KoboldCpp · KoboldCpp vs Lemonade · KoboldCpp vs llama.cpp · KoboldCpp vs LM Studio · KoboldCpp vs LocalAI · KoboldCpp vs MLX LM · KoboldCpp vs Ollama · KoboldCpp vs Open WebUI · KoboldCpp vs screenpipe · KoboldCpp vs TextGen · KoboldCpp vs Underdog
Machine-readable
| Similar tool | Grade | Score | Shared capabilities | x402 |
|---|---|---|---|---|
| LocalAI Ettore Di Giacinto and the LocalAI team | B | 68 | inference.local inference.open-weights agent.mcp-client embed.text speech.stt speech.tts image.generate | no |
| Lemonade AMD and the Lemonade community | B | 63.8 | inference.local inference.open-weights embed.text speech.stt speech.tts image.generate | no |
| DeepInfra Deep Infra Inc. | B | 63 | inference.open-weights embed.text image.generate speech.stt speech.tts | no |
| TextGen oobabooga | E | 45.1 | inference.local inference.open-weights agent.mcp-client embed.text image.generate | no |
| Foundry Local Microsoft | C | 60.5 | inference.local inference.open-weights embed.text speech.stt | no |
| llama.cpp ggml.ai (Hugging Face) | C | 60.2 | inference.local inference.open-weights embed.text agent.mcp-client | no |
Machine-readable
- JSON
/api/v1/tools/koboldcpp.json· historyhistory.json· badge/badges/koboldcpp.svg· changes feed/feeds/tools/koboldcpp.xml - Markdown
/tools/koboldcpp.md· slim/tools/koboldcpp.min.md(or sendAccept: text/markdown) - Fix list
/fixes/koboldcpp.md·/fixes/koboldcpp.json - From a terminal
anchor tool koboldcpp --md(the CLI) · over MCPget_tool {"slug": "koboldcpp"}at/mcp, no key - Directory index
/api/v1/tools.json· site index/llms.txt
Verify this listing
For the vendorIs this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.
-
Add the badge or a link
On a light page On a dark page <a href="https://www.anchorterminal.com/tools/koboldcpp"><img src="https://www.anchorterminal.com/badges/koboldcpp.svg" alt="KoboldCpp on Anchor Terminal" height="20"></a>[](https://www.anchorterminal.com/tools/koboldcpp)<a href="https://www.anchorterminal.com/tools/koboldcpp">KoboldCpp on Anchor Terminal</a>It counts on a page on koboldcpp.net or one of its subdomains, or the README of github.com/LostRuins/koboldcpp.
-
Tell us where it is
We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.
Agents send the same to POST /api/v1/verify as {"slug": "koboldcpp", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.


