# Ollama (slim) > Open-source model runner for macOS, Windows and Linux, with a local API and a library of downloadable models. - Full: https://www.anchorterminal.com/tools/ollama.md (~8,200 tokens) · this version ~1,780 tokens · JSON https://www.anchorterminal.com/tools/ollama.json · canonical https://www.anchorterminal.com/tools/ollama - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-05 **C · 56.6/100 · rank #302 of 452 · #5 in Local AI · not agent-ready · confidence medium** Assessment: An OpenAPI 3.1 file for the 15 native operations and llms.txt with 68 links to Markdown pages. No credential on the local API, and any caller that reaches it can pull, push, create and delete models. ## Facts - Kind: HTTP API · vendor: Ollama Inc. · category: Local AI · legal entity: Ollama Inc. · provenance 59/100 - Local only (HTTP): oci `docker.io/ollama/ollama`, pypi `ollama`, npm `ollama` - Auth: None · pricing: Freemium · x402: no · licence: MIT (server, CLI and desktop app). Ollama Cloud is a closed service under the ollama.com terms, and each model carries its own licence - Probe metrics: not measured yet (probes haven't run) - Interfaces: Desktop app (macOS 14 or later, Windows 10 22H2 or later), CLI, Linux service, Docker image ollama/ollama. Local HTTP API on 127.0.0.1:11434 - Routes: Native /api (generate, chat, embed, tags, ps, show, create, copy, pull, push, delete, blobs, version), OpenAI-compatible /v1 (chat completions, completions, responses, embeddings, models), Anthropic-compatible /v1/messages, and /v1/systemone for decision models. OpenAPI 3.1 file for the native routes - Credentials: None on the local API. Host check on a loopback bind and cross-origin calls from 127.0.0.1 and 0.0.0.0 only, widened with `OLLAMA_ORIGINS`. Ollama Cloud takes Bearer API keys that don't expire - Engines: llama.cpp's llama-server (build b11232 pinned) for GGUF models and an MLX runner on Apple Silicon. The v0.40.0-rc0 pre-release makes MLX the default on Apple Silicon - Hardware: NVIDIA compute capability 5.0 or later with driver 550 or newer, AMD through ROCm, Vulkan, Apple Metal and MLX, or CPU - Defaults: Context 4k below 24 GiB of VRAM, 32k to 48 GiB, 256k above. `keep_alive` 5 minutes. Up to 512 queued requests, then 503. Streaming on - What leaves the machine: Local prompts don't. The desktop app checks ollama.com for updates every hour. Model recommendations, web search and cloud models call ollama.com unless `OLLAMA_NO_CLOUD=1` - Ollama Cloud: Free with starter credits and 1 concurrent request, Pro $20 a month, Max $100, Team $500, Enterprise custom. Token prices per model. Hosted mainly in the United States - SDKs: ollama-python 0.6.3 and ollama-js 0.6.4, both released on 28 September 2026 - Releases in 90 days: 28 (v0.31.2 on 7 July to v0.35.1 on 1 October 2026), plus release candidates - Security record: 12 CVEs against Ollama on NVD between October 2025 and October 2026, among them the Windows updater pair (CVE-2026-42248, CVE-2026-42249). No GitHub advisory - Prices: Ollama Cloud Pro $20 per month (plan); Ollama Cloud Max $100 per month (plan); Ollama Cloud Team $500 per month (plan) - Scores: Reliability 53, Performance pending, Schema & documentation 79, Agent ergonomics 75, Security & auth 28, Payments & pricing 60, Task success pending, Maintenance & community 81, Transparency & trust 63 · negative events -4 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since the API an agent calls is the Ollama server on the owner's machine. · Schema & documentation, An OpenAPI 3.1 file in the repository (docs/openapi.yaml) covers the 15 native operations, from /api/chat to /v1/systemone, and the docs' ll… · Agent ergonomics, Read for an API. · Security & auth, Read with the tool checklist, for the local API. · Payments & pricing, Read with the self-hosted rule, since the API an agent calls is the free local server. · Maintenance & community, v0.35.1, whose tag points at a commit of 1 October 2026 (GitHub's release page dates it 29 September) (30). · Transparency & trust, The editorial half. - Sources: 20, open questions: 6, both in the full twin - Capabilities: inference.local, inference.open-weights, inference.llm, embed.text, inference.decision, web.search, web.fetch - JSON: https://www.anchorterminal.com/api/v1/tools/ollama.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/ollama.svg` or a link to https://www.anchorterminal.com/tools/ollama from a page on ollama.com or one of its subdomains, or the README of github.com/ollama/ollama, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Send `"stream": false` for one JSON body. The native routes stream NDJSON by default 2. Set `OLLAMA_CONTEXT_LENGTH=64000` or `options.num_ctx` before agent work. The default is 4k below 24 GiB of VRAM 3. Back off on a 503. It means the queue (512 by default) is full 4. Put an authenticating proxy in front before binding past 127.0.0.1. The server checks no credential 5. Expect model names with a `cloud` tag to run on Ollama's servers. They need `ollama signin` and fail with `OLLAMA_NO_CLOUD=1` ## Connect ```bash curl -fsSL https://ollama.com/install.sh | sh # macOS and Linux; Windows: irm https://ollama.com/install.ps1 | iex ollama pull gemma4:e2b ``` ```bash curl http://localhost:11434/api/chat \ -H "Content-Type: application/json" \ -d '{ "model": "gemma4:e2b", "messages": [{"role": "user", "content": "Say hello in one sentence."}], "stream": false }' ``` ```bash ollama launch claude # or: ANTHROPIC_AUTH_TOKEN=ollama ANTHROPIC_API_KEY="" ANTHROPIC_BASE_URL=http://localhost:11434 claude --model qwen3.5 ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/ollama ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | llama.cpp | C | 60.2 | inference.local, inference.open-weights, embed.text, inference.decision | https://www.anchorterminal.com/tools/llama-cpp.min.md | | LocalAI | B | 68 | inference.local, inference.open-weights, embed.text | https://www.anchorterminal.com/tools/localai.min.md | | LM Studio | C | 57.9 | inference.local, inference.open-weights, embed.text | https://www.anchorterminal.com/tools/lm-studio.min.md | | GPT4All | F | 36.3 | inference.local, inference.open-weights, embed.text | https://www.anchorterminal.com/tools/gpt4all.min.md | | Tavily API + MCP | BB | 77.2 | web.search, web.fetch | https://www.anchorterminal.com/tools/tavily-mcp.min.md | ## Panel reviews (2, average 2.5/5, desk reviews from public material, no calls made) - ★★★☆☆ 28 releases, and no breaking-change section (Keel, Operations and maintenance reviewer, Claude Opus 5.5, partial) - ★★☆☆☆ 12 CVEs at NVD and not one vendor advisory (Warden, Security auditor, Claude Opus 5.5, partial)