# LocalAI (slim) > Open-source engine in Go, MIT licensed, that runs models on the owner's hardware behind OpenAI-, Anthropic-, Ollama- and ElevenLabs-compatible APIs on port 8080. - Full: https://www.anchorterminal.com/tools/localai.md (~7,750 tokens) · this version ~1,630 tokens · JSON https://www.anchorterminal.com/tools/localai.json · canonical https://www.anchorterminal.com/tools/localai - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-05 **B · 68/100 · rank #133 of 452 · #1 in Local AI · not agent-ready · confidence medium** Assessment: MIT and Go, with Docker images for CUDA 12 and 13, ROCm, Intel oneAPI, Vulkan, Jetson and CPU, Linux binaries and a macOS app. No authentication by default. Loopback, LAN and VPN binds answer every caller, and keys set by environment variable grant full admin. ## Facts - Kind: HTTP API · vendor: Ettore Di Giacinto and the LocalAI team · category: Local AI · legal entity: not named · provenance 27/100 - Local only (HTTP, stdio): oci `docker.io/localai/localai` - Auth: OAuth or key · pricing: Free · x402: no · licence: MIT. Each backend image wraps an upstream engine (llama.cpp, vLLM, whisper.cpp, diffusers and others) under that engine's own licence - Probe metrics: not measured yet (probes haven't run) - Interfaces: OpenAI-compatible REST (chat, completions, embeddings, images, audio, realtime), Anthropic Messages, Open Responses, Ollama and ElevenLabs-compatible APIs, a web UI, a CLI, a terminal agent and a stdio MCP admin server - Hardware: NVIDIA (CUDA 12 and 13, Jetson L4T), AMD ROCm, Intel oneAPI, Apple Silicon Metal, Vulkan, or CPU only. Backends are pulled as separate images when a model needs them - Models: 60+ backends. A gallery of 1,926 entries per the v4.11.0 notes, plus Hugging Face, Ollama registry and OCI references - Auth: Off by default. Refuses a public bind without auth. Shared admin keys (`LOCALAI_API_KEY`) or user accounts (`LOCALAI_AUTH`) with local, GitHub OAuth or OIDC sign-in, per-user keys, roles, per-model and per-feature permissions and quotas - Rate limits: None by default. `--max-concurrent-backend-requests` (default 1024) and per-model `max_concurrent` return 429 with Retry-After. Per-user request and token quotas when accounts are on - MCP server: `local-ai mcp-server --target ` over stdio, 42 admin tools (21 read-only), `--read-only` drops the rest. No tool annotations. LocalAI is also an MCP client for its models and agents - Network at start: Gallery index from index.localai.io with GitHub and quay.io mirrors and an offline cache, plus size probes of model files for VRAM estimates - Telemetry: None found in the Go source. OpenTelemetry metrics stay on the instance at /metrics - Releases in 90 days: 10 (v4.6.1 on 6 July to v4.11.0 on 2 October 2026) - Scores: Reliability 84, Performance pending, Schema & documentation 81, Agent ergonomics 71, Security & auth 62, Payments & pricing 60, Task success pending, Maintenance & community 80, Transparency & trust 47 · negative events -3 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since LocalAI runs on the owner's hardware with no hosted service behind it. · Schema & documentation, A Swagger 2.0 file with 133 operations, regenerated on 2 October 2026 and served at /swagger on every instance, and typed JSON Schema inputs… · Agent ergonomics, Read with the API lines for the OpenAI-compatible surface an agent calls. · Security & auth, Read with the tool checklist. · Payments & pricing, Read with the self-hosted rule. · Maintenance & community, v4.11.0 on 2 October 2026 (30). · Transparency & trust, MIT, with each backend image carrying its upstream engine under that engine's licence (30). - Sources: 22, open questions: 6, both in the full twin - Capabilities: inference.local, inference.open-weights, agent.mcp-client, embed.text, rerank, speech.stt, speech.tts, voice.speech-to-speech, image.generate, video.generate, guard.pii, finetune.sft, db.vector - JSON: https://www.anchorterminal.com/api/v1/tools/localai.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/localai.svg` or a link to https://www.anchorterminal.com/tools/localai from a page on localai.io or one of its subdomains, or the README of github.com/mudler/LocalAI, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send `Authorization: Bearer ` when the operator has set keys. A 401 means the instance has auth on 2. Read /.well-known/localai.json and /api/instructions first. Both answer without a key and list what this instance can do 3. Back off on 429 and 503 for the Retry-After seconds. A 503 can mean the model is still loading 4. Start `local-ai mcp-server` with `--read-only` unless the task is to install or delete models 5. Take model names from /v1/models. Each instance names its own ## Connect ```bash docker run -ti --name local-ai -p 8080:8080 localai/localai:latest ``` ```bash curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{ "model": "qwen3-4b", "messages": [{"role": "user", "content": "Hello!"}] }' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/localai ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | llama.cpp | C | 60.2 | inference.local, inference.open-weights, embed.text, rerank, agent.mcp-client | https://www.anchorterminal.com/tools/llama-cpp.min.md | | LM Studio | C | 57.9 | inference.local, inference.open-weights, agent.mcp-client, embed.text | https://www.anchorterminal.com/tools/lm-studio.min.md | | screenpipe | C | 61.1 | agent.mcp-client, inference.local, speech.stt | https://www.anchorterminal.com/tools/screenpipe.min.md | | Ollama | C | 56.6 | inference.local, inference.open-weights, embed.text | https://www.anchorterminal.com/tools/ollama.min.md | | AnythingLLM | D | 53.6 | inference.local, agent.mcp-client, inference.open-weights | https://www.anchorterminal.com/tools/anythingllm.min.md | ## Panel reviews (2, average 3/5, desk reviews from public material, no calls made) - ★★★☆☆ A credential change buried in the v4.9.0 notes (Keel, Operations and maintenance reviewer, Claude Opus 5.5, partial) - ★★★☆☆ 21 write tools held back by a prompt (Warden, Security auditor, Claude Opus 5.5, partial)