# Lemonade (slim) > Open-source local AI server from AMD and community contributors. It runs text, speech and image models on the owner's CPU, GPU or NPU behind OpenAI-, Anthropic- and Ollama-compatible APIs and an MCP endpoint on port 13305. - Full: https://www.anchorterminal.com/tools/lemonade.md (~7,300 tokens) · this version ~1,730 tokens · JSON https://www.anchorterminal.com/tools/lemonade.json · canonical https://www.anchorterminal.com/tools/lemonade - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **B · 63.8/100 · rank #336 of 842 · #2 in Local AI · not agent-ready · confidence medium** Assessment: Apache 2.0, with weekly releases, installers for Windows, macOS and five Linux routes, and a six-tool MCP endpoint whose descriptions steer a caller away from accidental multi-gigabyte downloads. Authentication is off by default, the repository has no security policy, and the docs tell WebSocket clients to pass the key in the URL. ## Facts - Kind: HTTP API · vendor: AMD and the Lemonade community · category: Local AI · legal entity: Advanced Micro Devices, Inc. · provenance 57/100 - Local only (HTTP): oci `ghcr.io/lemonade-sdk/lemonade-server` - Auth: API key · pricing: Free · x402: no · licence: Apache 2.0. Each backend (llama.cpp, whisper.cpp, stable-diffusion.cpp, FastFlowLM and others) is downloaded separately under its own licence - Probe metrics: not measured yet (probes haven't run) - Interfaces: OpenAI-compatible REST under `/v1` and `/api/v1` (chat, completions, embeddings, responses, audio, images, realtime WebSocket), Anthropic Messages at `POST /v1/messages`, Ollama routes under `/api`, llama.cpp routes, Lemonade's own management API, an MCP endpoint at `POST /mcp`, a CLI (`lemonade`), a web UI and a tray app - MCP tools: 6 over Streamable HTTP (`lemonade_list_models`, `lemonade_chat`, `lemonade_transcribe_audio`, `lemonade_generate_image`, `lemonade_omni`, `lemonade_docs`). No annotations, no streaming, no embeddings or text-to-speech tool - Hardware: llama.cpp on CPU, Vulkan, ROCm, CUDA (Turing or newer) and Apple Metal. AMD XDNA2 NPUs through FastFlowLM (Windows, Linux) and Ryzen AI (Windows). vLLM on Strix Halo is marked experimental - Models: GGUF, FLM and ONNX language models, whisper.cpp and Moonshine speech-to-text, Kokoro text-to-speech, stable-diffusion.cpp images, pulled from Hugging Face or ModelScope - Install: Windows MSI and winget (`AMD.LemonadeServer`), macOS pkg, Ubuntu PPA and Snap, Debian deb, Fedora rpm, an Arch package, and a container at `ghcr.io/lemonade-sdk/lemonade-server` - Auth: Off by default. `LEMONADE_API_KEY` gates the regular API and `LEMONADE_ADMIN_API_KEY` gates `/internal/*`, both as bearer tokens set by environment variable - Rate limits: None documented. `max_loaded_models` (default 1 per model type) evicts the least recently used model - Network by default: Model and backend downloads from Hugging Face, ModelScope and GitHub, a start-up update check of downloaded models, and UDP broadcast for discovery. `offline: true` and `no_fetch_executables: true` block downloads - Telemetry: Off by default. When enabled, OTLP traces go to the operator's own collector, with `hide_inputs`, `hide_outputs` and `hide_thinking` to redact text - Releases in 90 days: At least 6 stable (v11.7.0 on 19 August to v2026.41.1 on 7 October 2026), plus weekly candidates - Scores: Reliability 75, Performance pending, Schema & documentation 70, Agent ergonomics 70, Security & auth 36, Payments & pricing 60, Task success pending, Maintenance & community 76, Transparency & trust 64 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since Lemonade runs on the owner's hardware with no hosted service behind it. · Schema & documentation, We found no OpenAPI or Swagger file in the repository. · Agent ergonomics, Read with the MCP lines for `/mcp` and the API lines for the OpenAI-compatible routes. · Security & auth, Read with the tool checklist. · Payments & pricing, Read with the self-hosted rule. · Maintenance & community, Stable release v2026.41.1 on 7 October 2026 (30). · Transparency & trust, Apache 2.0, with backends downloaded separately under their own licences (30). - Sources: 27, open questions: 8, both in the full twin - Capabilities: inference.local, inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate - JSON: https://www.anchorterminal.com/api/v1/tools/lemonade.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/lemonade.svg` or a link to https://www.anchorterminal.com/tools/lemonade from a page on lemonade-server.ai or one of its subdomains, or the README of github.com/lemonade-sdk/lemonade, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Call `lemonade_list_models` (or `GET /v1/models`) before naming a model. A wrong name with `allow_download: true` can start a multi-gigabyte download 2. Send `Authorization: Bearer ` when the operator has set `LEMONADE_API_KEY`. `/internal/*` needs the admin key when one is set 3. Use `POST /v1/chat/completions` for streamed tokens, embeddings and speech. The MCP endpoint ignores `stream` and has no embeddings or text-to-speech tool 4. Pass `output_dir` to `lemonade_generate_image` and `lemonade_omni` to get file paths in place of inline base64. Writes stay inside the MCP media sandbox 5. Read `GET /v1/docs` on the running server for the reference that matches its version. Weekly releases change behaviour ## Connect ```bash winget install --id AMD.LemonadeServer -e ``` ```bash curl http://localhost:13305/api/v1/chat/completions -H "Content-Type: application/json" -d '{"model": "your-model-name", "messages": [{"role": "user", "content": "Hello"}]}' ``` ```bash lemonade launch claude ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/lemonade ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.local, inference.open-weights, embed.text, rerank, speech.stt, speech.tts, image.generate | https://www.anchorterminal.com/tools/localai.min.md | | DeepInfra | B | 63 | inference.open-weights, embed.text, rerank, image.generate, speech.stt, speech.tts | https://www.anchorterminal.com/tools/deepinfra.min.md | | KoboldCpp | C | 60.5 | inference.local, inference.open-weights, embed.text, image.generate, speech.stt, speech.tts | https://www.anchorterminal.com/tools/koboldcpp.min.md | | Docker Model Runner | C | 57.1 | inference.local, inference.open-weights, embed.text, rerank, image.generate | https://www.anchorterminal.com/tools/docker-model-runner.min.md | | Foundry Local | C | 60.5 | inference.local, inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/foundry-local.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)