# Foundry Local (slim) > Microsoft's on-device model runtime, built on ONNX Runtime. Applications embed it through SDKs for C#, JavaScript, Python and Rust, and it can start an optional OpenAI-compatible server on localhost. A preview CLI is also available. - Full: https://www.anchorterminal.com/tools/foundry-local.md (~8,200 tokens) · this version ~1,780 tokens · JSON https://www.anchorterminal.com/tools/foundry-local.json · canonical https://www.anchorterminal.com/tools/foundry-local - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 60.5/100 · rank #461 of 842 · #4 in Local AI · not agent-ready · confidence medium** Assessment: The SDK and its native runtime are MIT, at 2.1.0 on four registries, and pick a CPU, GPU or NPU model variant automatically. The optional local server has no credential, its current routes aren't in the published REST reference, and telemetry is on by default with an opt-out. ## Facts - Kind: SDK + MCP · vendor: Microsoft · category: Local AI · legal entity: Microsoft Corporation · provenance 80/100 - Local only (HTTP): npm `foundry-local-sdk`, pypi `foundry-local-sdk`, nuget `Microsoft.AI.Foundry.Local`, cargo `foundry-local-sdk` - Auth: None · pricing: Free · x402: no · licence: MIT for the SDKs and the v2 native runtime. The CLI is closed source under Microsoft Software Licence Terms. Execution providers carry NVIDIA, Intel and Qualcomm licences, and each model carries its own - Probe metrics: not measured yet (probes haven't run) - Interfaces: SDKs for C#, JavaScript, Python and Rust, with C++, a C ABI and Java in the v2 source. Optional OpenAI-compatible server started by `start_web_service()` or `foundry server start`. `foundry` CLI in public preview - Local server: http://127.0.0.1 on a dynamic port by default. `/v1/chat/completions`, `/v1/responses` (stored, with `previous_response_id`), `/v1/embeddings`, `/v1/audio/transcriptions`, `/v1/models`, `/models/load/{name}`, `/models/unload/{name}`, `/models/loaded`, `/status`, `POST /shutdown`, per the v2 source - Credentials: None. The docs tell OpenAI clients to send any placeholder key and set Open WebUI's auth to None - Engine: ONNX Runtime 1.30.0 and ONNX Runtime GenAI 0.17.1 in v2. Models are ONNX, from the Foundry catalogue or registered by the owner (bring your own model, 2.1.0) - Hardware: CPU everywhere. WebGPU through Dawn on Windows, Linux and macOS (Metal on Apple silicon). NVIDIA CUDA on Windows and Linux. Intel OpenVINO, Qualcomm QNN and AMD Vitis AI on Windows - Runs on: Windows, macOS on Apple silicon and Linux, x64 and ARM64. Node.js 20 or later, Python 3.11 to 3.14, .NET 8 or later, Rust 1.70 or later - Models: A curated catalogue. The docs name GPT OSS, Qwen, DeepSeek, Mistral and Phi for chat and Whisper for transcription, plus embedding models. Each model has its own licence - What leaves the machine: Model and execution provider downloads, and telemetry events through Microsoft's 1DS SDK, on by default. The privacy file says prompts, outputs and audio are not collected - Releases in 90 days: 6 (CLI preview 0.10.2 on 14 July, v1.2.4, CLI preview 0.10.3, v2.0.1, v2.1.0, CLI preview 0.11.0 on 7 October 2026) - Packages: foundry-local-sdk 2.1.0 on npm (105,372 downloads in the week to 4 October 2026), PyPI (47,344 in the last week) and crates.io, and Microsoft.AI.Foundry.Local 2.1.0 on NuGet. CLI through winget and Homebrew - Issues: 76 open and 362 closed on 8 October 2026, with 32 open pull requests. No published security advisory and no CVE found at NVD - Scores: Reliability 68, Performance pending, Schema & documentation 53, Agent ergonomics 61, Security & auth 39, Payments & pricing 60, Task success pending, Maintenance & community 88, Transparency & trust 72 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since Foundry Local runs in the owner's application or as a local daemon. · Schema & documentation, Read with the API lines, for the local server and the SDKs. · Agent ergonomics, Read with the API lines. · Security & auth, Read with the tool checklist, for the local server. · Payments & pricing, Read with the self-hosted rule, since everything an agent calls runs on the owner's machine. · Maintenance & community, SDK 2.1.0 reached the registries on 29 September 2026 and CLI preview 0.11.0 followed on 7 October (30). · Transparency & trust, The editorial half. - Sources: 27, open questions: 7, both in the full twin - Capabilities: inference.local, inference.open-weights, embed.text, speech.stt - JSON: https://www.anchorterminal.com/api/v1/tools/foundry-local.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/foundry-local.svg` or a link to https://www.anchorterminal.com/tools/foundry-local from a page on microsoft.com or foundrylocal.ai or one of their subdomains, or the README of github.com/microsoft/Foundry-Local, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Read the server URL from `manager.urls[0]` or `foundry server status`. The port is dynamic unless the owner sets `web.urls` or `foundry server start --port` 2. Send the model ID that `GET /v1/models` returns, not the alias. The alias resolves to a hardware-specific variant 3. Check `supportsToolCalling` before sending tools. Support differs by variant, and issue #1183 reports Qwen tool calling failing on QNN 4. Set your own request timeout. Inference has no built-in one, and cancellation takes effect only after the current generation step 5. Ask the owner to set `ORT_TELEMETRY_DISABLED=1` or `disableNonessentialTelemetry` before the manager is created if telemetry must be off ## Connect ```bash pip install foundry-local-sdk # or: npm install foundry-local-sdk # CLI (preview): winget install Microsoft.FoundryLocal # macOS: brew tap microsoft/foundrylocal && brew install foundrylocal ``` ```bash foundry server start --port 39839 --idle-timeout 0 curl http://localhost:39839/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "", "messages": [{"role": "user", "content": "What is the golden ratio?"}]}' ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/foundry-local ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.local, inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/localai.min.md | | Lemonade | B | 63.8 | inference.local, inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/lemonade.min.md | | KoboldCpp | C | 60.5 | inference.local, inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/koboldcpp.min.md | | DeepInfra | B | 63 | inference.open-weights, embed.text, speech.stt | https://www.anchorterminal.com/tools/deepinfra.min.md | | llama.cpp | C | 60.2 | inference.local, inference.open-weights, embed.text | https://www.anchorterminal.com/tools/llama-cpp.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)