Head to head · Local inference · October 2026 research run

Foundry Local vs MLX LM

Foundry Local scores 60.5 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 6 of 7 scored categories. Both do local inference.

Which one, for what

Foundry Local C

Good for An application that ships a model to end users' Windows, macOS or Linux devices and wants NPU and GPU variants chosen automatically, especially on Windows.

Ahead on

  • Schema & documentation, 53 against 37
  • Agent ergonomics, 61 against 54
  • Security & auth, 39 against 32
  • Maintenance & community, 88 against 61
  • Transparency & trust, 72 against 66

Watch for

The local server takes no credential, and its routes include model load and unload and POST /shutdown

MLX LM D

Good for An owner with an Apple silicon Mac who wants MLX-format models, local fine-tuning and quantisation from Python or the command line, with a simple local chat completions server.

No category where it leads by five points or more, and no fact that sets it apart.

Watch for

mlx_lm.server has no API key or other credential option, and --allowed-origins defaults to *

Score by category

CategoryWeight this runFoundry LocalMLX LMEdge
Reliability16%206866Foundry Local +2
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25337Foundry Local +16
Agent ergonomics13%16.26154Foundry Local +7
Security & auth14%17.53932Foundry Local +7
Payments & pricing10%12.56060even
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88861Foundry Local +27
Transparency & trust7%8.87266Foundry Local +6
Negative events≤1500
Total60.5 · C52.2 · D

Facts side by side

FactFoundry LocalMLX LM
KindSDK + MCPHTTP API
VendorMicrosoftApple Inc.
Hosted endpointno (local only)no (local only)
TransportsHTTPHTTP
AuthNoneNone
PricingFreeFree
x402nono
LicenceMIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its ownMIT
Read-only variant documentednono
llms.txtnono
Last release2026-09-292026-10-01
Terms last updatedno document linkedno document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train models
Terms restrict automated access
Terms restrict benchmarking
Terms or service can change without notice
Arbitration or class-action waiver
Popularity2.6k stars, 105k npm/wk, 47k PyPI/wk7.3k stars, 140k PyPI/wk

Verdicts

Foundry Local

The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out.

MLX LM

MIT, with no telemetry found in the source, and the tests passed on the last eight pushes to main. mlx_lm.server has no API key option, answers any origin by default and loads whichever model a request names, and its own docs say it is not recommended for production.

Before you call either

Foundry Local

  1. Read the server URL from manager.urls[0] or foundry server status. The port is dynamic unless the owner sets web.urls or foundry server start --port
  2. Send the model ID that GET /v1/models returns, not the alias. The alias resolves to a hardware-specific variant
  3. Check supportsToolCalling before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN
  4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step
  5. Ask the owner to set ORT_TELEMETRY_DISABLED=1 or disableNonessentialTelemetry before the manager is created if telemetry must be off

MLX LM

  1. Keep mlx_lm.server on 127.0.0.1 and pass --allowed-origins with the origins you trust. There is no API key, and the default answers every origin
  2. Treat any caller as able to load any model. The model and adapters request fields accept any Hugging Face repository or local path
  3. Send max_tokens or max_completion_tokens when you need more than 512 tokens, the server default
  4. Read errors as {"error": "<text>"} with 400 for a bad field and 404 for a model that failed to load. They are not OpenAI error objects
  5. Poll GET /health before the first request. It answers 503 with unavailable when the generation thread has stopped

Questions

Which is better for AI agents, Foundry Local or MLX LM?

Foundry Local scores 60.5 (C) on agent readiness against MLX LM's 52.2 (D), and leads in 6 of 7 scored categories.

Can an agent call Foundry Local and MLX LM without installing anything?

No hosted endpoint is listed for Foundry Local. No hosted endpoint is listed for MLX LM.

Are Foundry Local and MLX LM open source?

Yes. Foundry Local is open source (MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own). MLX LM is open source (MIT).

Other comparisons with Foundry Local or MLX LM

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.