# vLLM (slim) > vLLM is an open-source inference and serving engine for open-weight language models. vllm serve runs an HTTP server with OpenAI-compatible, Anthropic Messages, embedding, reranking and transcription routes on the owner's own GPUs or CPUs. - Full: https://www.anchorterminal.com/tools/vllm.md (~7,650 tokens) · this version ~1,930 tokens · JSON https://www.anchorterminal.com/tools/vllm.json · canonical https://www.anchorterminal.com/tools/vllm - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 57.7/100 · rank #600 of 950 · #8 in Local AI · not agent-ready · confidence medium** Assessment: Apache-2.0 software with a release about every two weeks, each with notes that list breaking changes and security fixes. The optional API key covers only some path prefixes, so `/invocations` and control routes such as `/pause` answer without it, and at least 81 security advisories were published in the 12 months to 9 October 2026. ## Facts - Kind: HTTP API · vendor: vLLM project (PyTorch Foundation) · category: Local AI · legal entity: The Linux Foundation (vLLM is a PyTorch Foundation project) · provenance 53/100 - Local only (HTTP): pypi `vllm`, oci `vllm/vllm-openai` - Auth: None · pricing: Free · x402: no · licence: Apache-2.0 - Probe metrics: not measured yet (probes haven't run) - Interfaces: `vllm serve` HTTP server on port 8000, optional gRPC Inference and Control services (`--grpc-port`), the `vllm` Python library (`LLM`, `SamplingParams`), and Docker images `vllm/vllm-openai` for CUDA, ROCm, CPU and XPU - Routes: OpenAI `/v1/chat/completions`, `/v1/completions`, `/v1/responses`, `/v1/embeddings`, `/v1/audio/transcriptions`, `/v1/audio/translations`, `/v1/realtime` and `/v1/models`, Anthropic `/v1/messages` and `/v1/messages/count_tokens`, Cohere `/v2/embed` and `/v2/rerank`, `/v1/systemone`, `/pooling`, `/classify`, `/score`, `/tokenize`, `/detokenize`, `/health`, `/metrics`, and SageMaker `/invocations` - Credentials: None by default. `--api-key` (one or several) or `VLLM_API_KEY`, sent as `Authorization: Bearer`, checked only under `/v1`, `/v2`, `/inference` and `/cohere`. No scopes. gRPC has none - Network defaults: Binds every interface on port 8000 when `--host` is unset. CORS origins, methods and headers `*`, credentials off. TLS through `--ssl-keyfile` and `--ssl-certfile`. Inter-node traffic is unencrypted - Hardware: NVIDIA, AMD and Intel GPUs and x86, Arm and PowerPC CPUs per the README, with plugins for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend, Apple Silicon (vLLM-Metal) and others. Linux, Python 3.11 to 3.14 - Models: Hugging Face model repositories, more than 200 architectures per the README, including multimodal, embedding, reranking and speech recognition models. Downloads from Hugging Face or ModelScope - Agent tools: Tool calling with per-model parsers, structured outputs (JSON schema, choice, regex, grammar), reasoning parsers, and tool servers including MCP through the Responses API, off by default - Telemetry: Usage statistics to stats.vllm.ai, on by default. Off with `VLLM_NO_USAGE_STATS=1`, `DO_NOT_TRACK=1` or the file `~/.config/vllm/do_not_track` - Install: `uv pip install vllm --torch-backend=auto` or pip from PyPI, a ROCm wheel index at wheels.vllm.ai, and Docker images - Releases in 90 days: Eight stable releases from v0.25.1 (12 July 2026) to v0.31.0 (tagged 2 October 2026, notes published 5 October), plus release candidates - Governance: A PyTorch Foundation project under the Linux Foundation, contributed by UC Berkeley in July 2024. Eleven project leads form the technical steering committee - Security record: At least 100 published GitHub advisories, 81 in the 12 months to 9 October 2026 (2 critical, 19 high). A vulnerability management team, private reporting and CVEs - Scores: Reliability 62, Performance pending, Schema & documentation 68, Agent ergonomics 64, Security & auth 50, Payments & pricing 60, Task success pending, Maintenance & community 88, Transparency & trust 67 · negative events -6 · total over the 7 assessed categories - Why: Reliability, Read with the local-software lines, since vLLM runs on the owner's hardware with no hosted service. · Schema & documentation, The server is a FastAPI application with Pydantic request models, so a running instance serves its own OpenAPI document and Swagger page unl… · Agent ergonomics, Read for an API. · Security & auth, Read with the tool checklist. · Payments & pricing, Read with the self-hosted rule. · Maintenance & community, v0.31.0 was tagged on 2 October 2026 and its notes published on 5 October (30). · Transparency & trust, The editorial half. - Sources: 24, open questions: 8, both in the full twin - Capabilities: inference.local, inference.open-weights, embed.text, rerank, speech.stt, inference.decision, agent.mcp-client - JSON: https://www.anchorterminal.com/api/v1/tools/vllm.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/vllm.svg` or a link to https://www.anchorterminal.com/tools/vllm from a page on vllm.ai or one of its subdomains, or the README of github.com/vllm-project/vllm, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Put a reverse proxy that allowlists routes in front of the server. `--api-key` leaves `/invocations` and the control routes open 2. Pass `--host 127.0.0.1` for single-machine use. With no `--host` the server listens on every interface 3. Set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1` before starting if nothing should be sent to stats.vllm.ai 4. Start with `--enable-auto-tool-choice` and the `--tool-call-parser` for the model before sending tools. Tool calling is off without them 5. Send `max_tokens` on every request, and read the breaking changes section of the release notes before upgrading a minor version ## Connect ```bash uv pip install vllm --torch-backend=auto vllm serve Qwen/Qwen2.5-1.5B-Instruct # listens on port 8000 ``` ```bash curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen2.5-1.5B-Instruct", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Who won the world series in 2020?"} ] }' ``` ```bash ANTHROPIC_BASE_URL=http://localhost:8000 \ ANTHROPIC_API_KEY=dummy \ ANTHROPIC_AUTH_TOKEN=dummy \ ANTHROPIC_DEFAULT_OPUS_MODEL=my-model \ ANTHROPIC_DEFAULT_SONNET_MODEL=my-model \ ANTHROPIC_DEFAULT_HAIKU_MODEL=my-model \ claude ``` Full config and headless snippets are in the full page. Through letme (picks today, calling later): https://letme.dev/vllm ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | LocalAI | B | 68 | inference.local, inference.open-weights, agent.mcp-client, embed.text, rerank, speech.stt | https://www.anchorterminal.com/tools/localai.min.md | | llama.cpp | C | 60.2 | inference.local, inference.open-weights, embed.text, rerank, inference.decision, agent.mcp-client | https://www.anchorterminal.com/tools/llama-cpp.min.md | | Lemonade | B | 63.8 | inference.local, inference.open-weights, embed.text, rerank, speech.stt | https://www.anchorterminal.com/tools/lemonade.min.md | | KoboldCpp | C | 60.5 | inference.local, inference.open-weights, agent.mcp-client, embed.text, speech.stt | https://www.anchorterminal.com/tools/koboldcpp.min.md | | Cloudflare Workers AI | B | 68.3 | inference.open-weights, embed.text, rerank, speech.stt | https://www.anchorterminal.com/tools/cloudflare-workers-ai.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)