# Fix list: LocalGhost From Anchor Terminal's listing at https://www.anchorterminal.com/tools/localghost, the October 2026 research run, assessed 3 October 2026. Grade E, 45.8 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on LocalGhost: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Agent ergonomics, 2 out of 100, up to 15.9 more on the total Why it scored 2: Read with the tool lines, as for any platform, against the surface an agent can reach. There's none. The box answers every caller but its paired phone as if it were down, and the README and llms.txt say there's no API, MCP server or endpoint for agents in this release. `ghost-cli` is the box's service console, needs root or the daemons' user on an unlocked box and is documented as the operator's tool, so it isn't scored as an agent's interface. No tool definitions or responses an agent can load or size (0). No paging, filtering or output-size controls on any surface an agent can reach (0). Refused, unauthenticated and retired callers all get one 503 identical to an outage. connection.md, the README's section for agents and llms.txt say so, so an agent knows in advance, but nothing in the response can be acted on (2 of 20). No idempotency keys, annotations or retry guidance on any surface an agent can reach (0). No SDK and no agent-facing defaults (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 2. Schema & documentation, 44 out of 100, up to 9.1 more on the total Why it scored 44: Read with the API lines, since the listing's kind is platform, and scored on what a model can read and call. `ghost-cli` is the operator's console. Its own source calls it the box's service console, SECURITY.md lists it among the operator's tools, it needs root or the daemons' user on an unlocked box, and the README says nothing in this release opens the box to an agent, so its commands aren't graded as an agent's interface here or under Agent ergonomics. No OpenAPI, MCP schema or other contract for the phone's HTTP routes or the daemons' control sockets, and server/docs/INGEST_CONTRACT.md still describes Bearer device tokens and 401 answers that the shipped code doesn't use (0). llms.txt at www.localghost.ai, current to 0.0.3, plus Markdown READMEs and release notes in the repository (10). The README's section for agents and llms.txt say what the box is, that there's no API or MCP server and that the way to a person's data is through that person, which tells an agent when not to use it. Some documents a model would read are stale. internal/secd/README.md says the real backend has never run, and CONTRIBUTING.md says there's one release (11 of 20). No input surface is open to an agent, so there's nothing typed to grade (0). Worked operator commands run through the 555-line setup guide and the release notes, and connection.md documents that every refused call gets the same 503 as a down box (8 of 15). Semver tags, RELEASES.md, notes per release in server/releases/ that say what still works with older versions, and a changelog page whose entries carry dates and, for releases, the tag and commit (15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 3. Security & auth, 52 out of 100, up to 8.4 more on the total Why it scored 52: Read with the tool checklist, credential model first. Each phone presents a client certificate from the box's own CA (ECDSA P-256, valid ten years, no scopes), swapped since 2 October 2026 for a non-exportable Android Keystore key at the first unlock, plus a PIN-unlocked session token of at most 48 hours in a header, and one phone can be retired by its key since 0.0.2. No credential exists for an agent. The enrolment QR carries the first key in a `localghost://enroll` query string, but it goes screen to camera, is drawn only on an interactive terminal and is retired after the first unlock, so the query-string deduction doesn't apply (20). Each paired phone reaches the whole archive, with no read-only or scoped device. A fresh box runs nginx's header path, where ghost.secd believes an `X-Client-Cert` header any local process on the box can forge, until the operator runs `ghost-ctl edge-passthrough` (setup guide step 8d, flagged by privacy_check.sh), and the known gaps of 0.0.2 and 0.0.3 no longer list it. Destructive setup steps need typed confirmations (YES, WIPEIT, APPLY), setup is a dry run until `--apply`, a release on trial rolls back by itself, and the daemons run as a user with no login shell (7 of 20). The box's model reads news feeds and pages the phone fetched, and we found no injection guidance. The daemons own the control flow, the model calls no tools, and answers citing figures not in the facts are dropped (6 of 15). Each daemon logs to journald, rejected sessions are logged with the remote address, a fetch log keeps one row per outbound address and shows on Box Status, each chat answer keeps a record of the context it was given, and `ghost-cli ghost.secd devices` lists paired phones, but there's no audit trail of the phone's API calls (8 of 15). SECURITY.md with scope, PGP, a 72-hour acknowledgement, a 14-day response, 90-day disclosure and a commitment not to sue, a valid RFC 9116 security.txt to 2027-10-01, and release sums signed by the same key. No bounty, GitHub private reporting off and no advisories yet. The release binaries and CI use Go 1.25.4, which the Go vulnerability database (commit 5cd8418, 1 October 2026) lists as affected by 35 standard-library advisories fixed in 1.25.5 to 1.25.13, among them crypto/tls, crypto/x509 and net/http, which ghost.secd uses for the phone's TLS (11 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 4. Reliability, 59 out of 100, up to 8.2 more on the total Why it scored 59: Read with the local-software lines, because LocalGhost runs on its owner's hardware with no hosted service, and its status page covers only the website, the mirror and the releases. The v0.0.3 GitHub release (14 assets) holds a reproducible linux-amd64 server bundle, a signed APK and signed sums, and the runtimes are stated (Debian 13 on x86-64 with an NVIDIA GPU, a TPM 2.0 and a spare NVMe, Android 15 or later). The bundle is how a running box takes an update from the phone. A first box install is a source build (`make box`, then `tools/setup_llama.sh` builds llama.cpp from the mirror's source), with no package-manager route (12 of 20). CI went in on 2 October 2026. Of its 11 runs on main at 10:50 UTC on 3 October, the first four failed and the last seven passed, including the v0.0.3 tag commit c3726d3 and HEAD d4decab, running build, vet, every Go test (489 test functions in 170 files) against Postgres and a macOS vet. The app's JVM tests ran on push only in the first two runs, which failed, and now run only on pull requests or by hand, with no passing run, while README.md and CONTRIBUTING.md say CI runs them on every push (17 of 25). No open issues. The three outside items (pull request 1, issue 2 and pull request 3, from July) were merged or closed on 2 October, and the 0.0.3 weather change undid pull request 1's lock-screen fix until c3726d3 put it back 24 minutes later, before the tag (20 of 25). Semver tags, notes per release that say what keeps working across versions (an older app still works with a 0.0.3 box, and the app installs over the last one and keeps its enrolment) and a dated changelog. Three releases within 16 hours is a short record, and there's no breaking-change heading yet (10 of 15). 0.0.3, pre-1.0 (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 5. Payments & pricing, 60 out of 100, up to 5 more on the total Why it scored 60: Read with the self-hosted rule. No x402, MPP or L402, and nothing is sold (0). The software is free under MIT with no account and no card, so 20, 20 and 20 on the last three lines. Pre-built boxes are planned at the cost of parts and assembly plus a 30 per cent margin, and optional one-time daemon packages, none priced yet. The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 6. Maintenance & community, 71 out of 100, up to 2.5 more on the total Why it scored 71: wisp 0.0.3 on 3 October 2026 (30). Three releases in the last 90 days, 0.0.1 and 0.0.2 on 2 October and 0.0.3 on 3 October, all within 16 hours, so the count is met while the cadence is untested (20). The three outside items, pull request 1 (20 July 2026), issue 2 and pull request 3 (28 July), were merged, fixed in the tree or closed on 2 October, 66 to 74 days later, and none had a written reply at 10:50 UTC on 3 October. The 0.0.2 notes say what closed issue 2 and pull request 3, and CONTRIBUTING.md says an answer may take a week (8 of 25). No MCP server, registry entry or SDK, by design. Each release ships a signed APK and a server bundle built from its tag, on GitHub and the setup mirror, with a source archive matching the tag's 846 files, but no package registry (10 of 15). CI on every push since 2 October, passing for the server, with the app job never passing. Go 1.25.4 is pinned with 35 standard-library advisories fixed since, x/crypto sits at v0.31.0 of December 2024 against v0.57.0 and x/term at v0.27.0 against v0.46.0, and there's no Dependabot (3 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## 7. Transparency & trust, 77 out of 100, up to 2 more on the total Made of editorial 72, provenance 82. Why it scored 77: MIT for the server and the app, all of it public. The `LICENSE` line reads LocalGhostDao, which the terms page names as the copyright holder, while the company is LocalGhost.ai Ltd (30). The privacy page (updated 3 October 2026) covers the website, the mirror and the software, names LocalGhost.ai Ltd, says access logging is off (the public nginx config agrees) and lists what the box and the phone fetch. It still disagrees with the code, and other documents disagree with each other. The privacy page says the phone talks to the box and nothing else apart from a search and the daily mirror check, and that it sends no position, as do the README, the setup page and llms.txt, while the app sends currency amounts to api.frankfurter.app, who-is subjects to Wikipedia's summary API and new location fixes to Android's system Geocoder, a network call on some phones by its own code comment. The setup guide says the box reaches the network at setup only, and the essay What the Box Does Now (29 September, modified 3 October) says no daemon opens a connection, while ghost.tallyd polls seven exchanges every minute. Email to info@localghost.ai is kept to reply, with no retention period (14 of 30). No deprecation policy. Each release's notes say which older app or box still works, and CONTRIBUTING.md makes any schema drop a named, dated migration (8 of 20). No telemetry, analytics or crash-reporting library in the server or the app, and the phone's daily release check against the mirror is disclosed (20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Domain age: localghost.ai, registered 2025-12-17 (under a year) (0 of 15) ## Deductions Each comes off the total. A fixed and documented problem counts for less at the next check. - 2026-09-19 to 2026-10-03, unfixed at HEAD d4decab. The README ('One thing leaves the phone'), the privacy page ('The phone talks to your box and nothing else, apart from two things'), the setup page ('The phone sends one thing') and llms.txt ('no position leaves either device') describe less than the app sends. Since 19 September it passes each new location fix to Android's system Geocoder when the phone has moved 20 km or 12 hours have passed, a network call on some phones by its own code comment, and since 20 September it sends currency amounts to api.frankfurter.app and who-is subjects to Wikipedia's summary API, named only in the engineering journal, though the chat shows those two as sources when it uses them. Until 0.0.3 the same claim stood while the app sent the phone's position to Open-Meteo for a weather question naming no place. wisp 0.0.3 removed that call on 3 October and its notes, the privacy page, llms.txt and the changelog say so, so that part adds nothing. One misleading claim, still partly false, -3. https://github.com/LocalGhostDao/localghost/blob/main/app/android/app/src/main/java/com/localghost/app/sync/LocationLog.kt; https://github.com/LocalGhostDao/localghost/blob/main/app/android/app/src/main/java/com/localghost/app/net/WebSearch.kt; https://www.localghost.ai/privacy ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - Whether the 0.0.3 server bundle rebuilds byte for byte. The signature on SHA256SUMS verified, but Go 1.25.4 couldn't be fetched from our shell to rebuild it - Which job failed in CI runs 1 and 2, which ran both the server and the app jobs. Runs 3 and 4 ran only the server job, and run 5 passed once 136133c dropped a stale internal/hw test - Whether Android's system Geocoder sends the position over the network on a typical Android 15 phone. The app's own comment says it does on some phones - unchecked: whether the live web server runs the nginx config in the public web repository, with access logging off - The listing's licence read MIT for code and CC BY-SA 4.0 for hardware designs. No hardware designs are published and the README now says only Code MIT, so the licence field is set to MIT ## Weaknesses - No API, MCP server or SDK for agents, and every caller but the paired phone gets the same 503 as a box that is down - The app sends currency questions to Frankfurter, who-is questions to Wikipedia and location fixes to Android's geocoder, none of them named in the README, privacy page or llms.txt - Release binaries are built with Go 1.25.4, and the 35 standard-library advisories fixed in Go 1.25.5 to 1.25.13 include crypto/tls, crypto/x509 and net/http - Pre-1.0 at 0.0.3, three releases within 16 hours, CI one day old, and the app's 50 JVM test files have no passing CI run while the README and CONTRIBUTING.md say they run on every push - A fresh box trusts a device-certificate header any local process can forge until the operator runs `ghost-ctl edge-passthrough` ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Don't call a LocalGhost box. Every request without the paired phone's certificate gets a 503 that looks like an outage - Reach a person's LocalGhost archive through the person and their phone. Nothing in wisp 0.0.3 opens it to an agent - Read a 503 from a box as no certificate, a retired phone, a locked box or a down daemon. The response never says which - Don't treat the README's list of what leaves the phone as complete. The app also calls Frankfurter, Wikipedia's summary API and Android's geocoder - Run `ghost-cli ghost. commands` first if an operator hands you a root shell on an unlocked box. It lists command names only, takes key=value arguments and prints JSON ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.