# Vela 2.0 (slim) > Vela 2.0 is a family of four open-weight decision models from the vLLM Semantic Router project and KR Labs, released on 6 October 2026 under Apache-2.0 for routing, safety checks, personal-data spans and hallucination checks. - Full: https://www.anchorterminal.com/tools/vela.md (~6,850 tokens) · this version ~1,730 tokens · JSON https://www.anchorterminal.com/tools/vela.json · canonical https://www.anchorterminal.com/tools/vela - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-08 **B · 66.5/100 · rank #219 of 629 · #4 in Decision models · not agent-ready · confidence medium** Assessment: One self-hosted call answers routing, prompt-attack, personal-data and unsupported-claim questions with probabilities and character offsets, under Apache-2.0 with SHA-256 manifests. The models are days old and carry no Hub version tags, and the three larger sizes keep 74 to 89 per cent of their Decision 2.0 bases on the Jev Decision Index by the authors' figures. ## Facts - Kind: Model API · vendor: vLLM Semantic Router project and KR Labs · category: Decision models · legal entity: not named · provenance 27/100 - Local only (HTTP): pypi `vllm-sr` - Auth: None · pricing: Free · x402: no · licence: Apache-2.0 (weights, code and documentation). The 0.3B's tokeniser keeps the Gemma Terms of Use, and training data keeps its own licences - Probe metrics: not measured yet (probes haven't run) - Models: Vela-2.0-0.3B (307M encoder, from Decision-1.0-Kai), Vela-2.0-0.8B (756M), Vela-2.0-4B (4.2B) and Vela-2.0-9B (7.9B), the last three fine-tuned from Decision 2.0 Eos, Nox and Lux on Qwen3.5 backbones - Licence: Apache-2.0 for weights, code and documentation. The 0.3B's tokeniser keeps the Gemma Terms of Use. Training data isn't redistributed and keeps its own licences, some CC BY-SA - Question types: choice (2 to 255 options), noul (yes or no), score (2 to 10 ordered levels), set (any number of labels) and span (labelled character offsets), any mix in one request - Context: 8,192 tokens on the 0.3B and 16,384 a rendered sequence on the decoders. Span targets over 2,048 tokens are read in windows of up to 1,800 - Hardware: The 0.3B runs on a CPU or through ONNX. GPU with bf16 autocast is the evaluated setting for the decoders, about 17 GB of parameter memory for the 4B. fp16 isn't supported - Serving: `AutoModel.from_pretrained(..., trust_remote_code=True)` and `model.system_one(...)`, the bundled `vela2_serve.py` on `POST /v1/systemone`, or the router's model runtime on `POST /v1/decisions` (development channel) - Span heads: A router head trained on PII (17 types), unsupported claims and toxic spans, and on the decoders a broad head for open labels. The response names the head that answered - Errors: Bundled server 401, 413 and 422 with `detail`. Model runtime 400, 404, 413, 422, 429 and 503 with `{error: {code, message}}` - Hosted option: None - Languages: 17 listed on the cards, among them Arabic, Chinese, English, French, German, Hindi, Japanese, Korean and Spanish - Scores: Reliability 57, Performance pending, Schema & documentation 78, Agent ergonomics 79, Security & auth 60, Payments & pricing 60, Task success pending, Maintenance & community 84, Transparency & trust 48 · total over the 7 assessed categories - Why: Reliability, Scored on the local-package checklist, since Vela 2.0 is a model the owner runs. · Schema & documentation, Read for a model the owner serves. · Agent ergonomics, Read as an API an agent calls for a decision, as with Jev, Clef and Kev. · Security & auth, Read as software the owner runs. · Payments & pricing, Free Apache-2.0 software with nothing to buy, so 20 + 20 + 20 for pricing, free use and no sign-up. · Maintenance & community, Read for an open-weight model. · Transparency & trust, Apache-2.0 for the weights, code and documentation, on Apache-2.0 Decision 2.0 and Qwen3.5 bases, with MODIFICATIONS.md and ATTRIBUTIONS.md… - Sources: 18, open questions: 8, both in the full twin - Capabilities: inference.decision, guard.pii, guard.injection, guard.moderation, guard.self-host - JSON: https://www.anchorterminal.com/api/v1/tools/vela.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/vela.svg` or a link to https://www.anchorterminal.com/tools/vela from a page on vllm-sr.ai or one of its subdomains, or the README of github.com/vllm-project/semantic-router, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Pin a commit hash with `revision=` when loading from the Hub. The repositories have no tags and `main` has changed since launch 2. Send the served name in `model`, for example `vllm-sr/Vela-2.0-4B`. The bundled server answers 422 to any other name 3. Name span questions `pii`, `halu` or `toxic`, or set `"head": "router"`, to get the trained router head. Other labels go to the broad head 4. Keep input under 16,384 tokens a sequence (8,192 on the 0.3B). The bundled server answers 413 when the questions alone don't fit 5. Set `VELA2_API_KEY` before binding the bundled server beyond 127.0.0.1, and keep the model runtime on a trusted network ## Connect ```bash pip install torch "transformers>=5.17" safetensors tokenizers numpy fastapi uvicorn # from a local snapshot of vllm-sr/Vela-2.0-4B python vela2_serve.py --model . --device cuda --port 8001 ``` ```bash curl -s localhost:8001/v1/systemone -H 'content-type: application/json' \ -d '{"model":"vllm-sr/Vela-2.0-4B","state":"My card was charged twice and the parcel never arrived.","questions":{"issues":{"type":"set","instructions":"Which issues does the customer report?","criteria":{"billing":"payments, charges, refunds or invoices","shipping":"delivery of an order or a parcel","login":"signing in, passwords or account access"}}}}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | NVIDIA NeMo Guardrails | B | 68.4 | guard.injection, guard.pii, guard.moderation, guard.self-host | https://www.anchorterminal.com/tools/nemo-guardrails.min.md | | Lakera Guard (Check Point AI Guardrails) | C | 59.6 | guard.injection, guard.pii, guard.moderation, guard.self-host | https://www.anchorterminal.com/tools/lakera-guard.min.md | | Guardrails AI | D | 49.6 | guard.injection, guard.pii, guard.moderation, guard.self-host | https://www.anchorterminal.com/tools/guardrails-ai.min.md | | Google Cloud Model Armor | BB | 77.9 | guard.injection, guard.pii, guard.moderation | https://www.anchorterminal.com/tools/google-model-armor.min.md | | Amazon Bedrock Guardrails | BB | 74.8 | guard.injection, guard.pii, guard.moderation | https://www.anchorterminal.com/tools/amazon-bedrock-guardrails.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)