# MLX LM (slim) > Open-source Python package and command-line tools from Apple's MLX team for running, quantising and fine-tuning language models on Apple silicon. mlx_lm.server exposes a local HTTP API modelled on OpenAI's chat completions. - Full: https://www.anchorterminal.com/tools/mlx-lm.md (~6,400 tokens) · this version ~1,530 tokens · JSON https://www.anchorterminal.com/tools/mlx-lm.json · canonical https://www.anchorterminal.com/tools/mlx-lm - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **D · 52.2/100 · rank #657 of 842 · #12 in Local AI · not agent-ready · confidence medium** Assessment: MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. `mlx_lm.server` has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production. ## Facts - Kind: HTTP API · vendor: Apple Inc. · category: Local AI · legal entity: Apple Inc. · provenance 67/100 - Local only (HTTP): pypi `mlx-lm` - Auth: None · pricing: Free · x402: no · licence: MIT - Probe metrics: not measured yet (probes haven't run) - Interfaces: `mlx_lm.server` HTTP API, the Python API (`load`, `generate`, `stream_generate`, `convert`), and command-line tools including `mlx_lm.generate`, `mlx_lm.chat`, `mlx_lm.lora`, `mlx_lm.convert`, `mlx_lm.fuse`, `mlx_lm.evaluate`, `mlx_lm.cache_prompt` and `mlx_lm.manage` - Routes: POST `/v1/chat/completions` (also `/chat/completions`) and `/v1/completions`, GET `/v1/models` and `/health` - Credentials: None, and no flag to add one - Network defaults: Binds 127.0.0.1:8080. `--allowed-origins` defaults to `*`. Models download from Hugging Face on first use, or from ModelScope when `MLXLM_USE_MODELSCOPE` is true - Request defaults: `max_tokens` 512, `temperature` 0.0, `top_p` 1.0. Only `messages` is required - Models: MLX-format models from Hugging Face, with 4-bit and other quantised builds in the mlx-community organisation. The default for `mlx_lm.generate` and `mlx_lm.chat` is `mlx-community/Llama-3.2-3B-Instruct-4bit` - Tool calling: The chat route accepts `tools` and parses tool calls with one of 13 model-specific parsers. It does not run them, and `SERVER.md` does not document the field - Fine-tuning: Low-rank (LoRA) and full fine-tuning with `mlx_lm.lora`, including on quantised models, and adapters fused with `mlx_lm.fuse` - Platforms: Python 3.11 to 3.13. macOS on Apple silicon, with Linux classifiers and `cuda12`, `cuda13` and `cpu` extras. Wired memory for large models needs macOS 15 or later - Install: `pip install mlx-lm` or `conda install -c conda-forge mlx-lm` - Releases: 97 versions on PyPI since 0.0.1 on 12 January 2024. 0.32.0 on 1 October 2026, the one before 0.31.3 on 22 April 2026 - Activity: 127 commits from 82 authors on main in the 90 days to 8 October 2026. 151 open issues and 74 open pull requests - Security record: No published GitHub advisory. Reports go through GitHub private vulnerability reporting - Scores: Reliability 66, Performance pending, Schema & documentation 37, Agent ergonomics 54, Security & auth 32, Payments & pricing 60, Task success pending, Maintenance & community 61, Transparency & trust 66 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since the package and its server run on the owner's machine. · Schema & documentation, Read for the HTTP server, the surface an agent calls. · Agent ergonomics, Read for an API. · Security & auth, Read with the tool checklist. · Payments & pricing, Read with the self-hosted rule. · Maintenance & community, PyPI release 0.32.0 on 1 October 2026 (30). · Transparency & trust, The editorial half. - Sources: 19, open questions: 6, both in the full twin - Capabilities: inference.local, inference.open-weights - JSON: https://www.anchorterminal.com/api/v1/tools/mlx-lm.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/mlx-lm.svg` or a link to https://www.anchorterminal.com/tools/mlx-lm from a page on apple.com or one of its subdomains, or the README of github.com/ml-explore/mlx-lm, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Keep `mlx_lm.server` on 127.0.0.1 and pass `--allowed-origins` with the origins you trust. There is no API key, and the default answers every origin 2. Treat any caller as able to load any model. The `model` and `adapters` request fields accept any Hugging Face repository or local path 3. Send `max_tokens` or `max_completion_tokens` when you need more than 512 tokens, the server default 4. Read errors as `{"error": ""}` with 400 for a bad field and 404 for a model that failed to load. They are not OpenAI error objects 5. Poll `GET /health` before the first request. It answers 503 with `unavailable` when the generation thread has stopped ## Connect ```bash pip install mlx-lm mlx_lm.server --model mlx-community/Mistral-7B-Instruct-v0.3-4bit # listens on 127.0.0.1:8080 ``` ```bash curl localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "messages": [{"role": "user", "content": "Say this is a test!"}], "temperature": 0.7 }' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/mlx-lm ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.local, inference.open-weights | https://www.anchorterminal.com/tools/localai.min.md | | Lemonade | B | 63.8 | inference.local, inference.open-weights | https://www.anchorterminal.com/tools/lemonade.min.md | | Foundry Local | C | 60.5 | inference.local, inference.open-weights | https://www.anchorterminal.com/tools/foundry-local.min.md | | KoboldCpp | C | 60.5 | inference.local, inference.open-weights | https://www.anchorterminal.com/tools/koboldcpp.min.md | | llama.cpp | C | 60.2 | inference.local, inference.open-weights | https://www.anchorterminal.com/tools/llama-cpp.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)