# Granite Guardian (slim) > Granite Guardian is IBM's family of open-weight judge models. The current 8-billion-parameter release answers yes or no on whether a prompt, response, retrieved context or function call meets a built-in or custom criterion, and the owner runs it. - Full: https://www.anchorterminal.com/tools/granite-guardian.md (~7,150 tokens) · this version ~1,680 tokens · JSON https://www.anchorterminal.com/tools/granite-guardian.json · canonical https://www.anchorterminal.com/tools/granite-guardian - Index: https://www.anchorterminal.com/llms.txt · API: https://www.anchorterminal.com/api/v1/index.json · Updated: 2026-10-09 **C · 60.1/100 · rank #478 of 842 · #10 in Guardrails & safety filters · not agent-ready · confidence medium** Assessment: One ungated Apache 2.0 model judges harm, jailbreaks, RAG groundedness, function-call errors and custom criteria, with signed weights and published evaluation code. It is trained and tested on English only, each call checks one criterion, the 4.1 prompt format differs from 3.x, and IBM's watsonx.ai lists only the deprecated 3.0 model. ## Facts - Kind: Model API · vendor: IBM · category: Guardrails & safety filters · legal entity: International Business Machines Corporation · provenance 73/100 - Local only (HTTP) - Auth: None · pricing: Free · x402: no · licence: Apache 2.0 for the weights and the repository - Probe metrics: not measured yet (probes haven't run) - Model: `ibm-granite/granite-guardian-4.1-8b`, 8,380,551,168 parameters in BF16 safetensors, fine-tuned from `ibm-granite/granite-4.1-8b`. `config.json` gives 131,072 positions - Licence: Apache 2.0. IBM asks that derivatives put the maker's name before Granite and says the name is a registered trademark (https://www.ibm.com/granite/docs/model-standards/naming-guidance) - Access: Not gated on Hugging Face. Also `ollama run granite4.1-guardian:8b` and IBM's GGUF builds - Built-in criteria: Harm, social bias, jailbreaking, violence, profanity, sexual content, unethical behaviour. Context relevance, groundedness and answer relevance for RAG. Function-calling hallucination for agents - Custom criteria: Any rule in natural language in the `### Criteria:` section. IBM calls this bring your own criteria and says such criteria need testing - Input: A conversation whose last user message is a `` block with a think or no-think instruction, the criterion and a scoring schema. Documents go in `documents=` and tool schemas in `available_tools=` - Output: `yes` or `no`. Think mode puts a reasoning trace in `` tags first - Runtimes: `transformers` and vLLM in the card, Ollama and GGUF from IBM. The evaluation code asks for `vllm>=0.15.1` and `transformers>=4.52` and says an 8B model fits on one 80 GB GPU - IBM's figures: Version 4.1 without thinking. Safety F1 0.79 across ten sets, LLM-AggreFact balanced accuracy 0.760, FC Reward Bench 0.79, IFEval multi-constraint 0.844 - Languages: English only, per the model card - Integrity: `model.sig` is a sigstore signature. IBM documents verification with `model-signing` and calls model signing experimental (https://www.ibm.com/granite/docs/model-standards/signature-verification) - Hosted by IBM: watsonx.ai lists `ibm/granite-guardian-3-8b` only, marked deprecated, at $0.0002 per 1,000 tokens. Version 4.1 is not listed - Support: Issues at github.com/ibm-granite/granite-guardian and the Hugging Face community tab. Security reports go to IBM PSIRT - Scores: Reliability 56, Performance pending, Schema & documentation 60, Agent ergonomics 69, Security & auth 58, Payments & pricing 60, Task success pending, Maintenance & community 47, Transparency & trust 70 · total over the 7 assessed categories - Why: Reliability, Scored on the local-software lines, since the owner runs the model. · Schema & documentation, Read for a model the owner serves. · Agent ergonomics, Read as a classifier an agent calls. · Security & auth, Read as software the owner runs. · Payments & pricing, Self-hosted rule. · Maintenance & community, Read for an open-weight model. · Transparency & trust, Weights and repository are under the Apache 2.0 licence with no gate and no use policy attached (30 of 30). - Sources: 19, open questions: 8, both in the full twin - Capabilities: guard.moderation, guard.policy, guard.injection, guard.self-host - JSON: https://www.anchorterminal.com/api/v1/tools/granite-guardian.json - Verify (for the vendor): the badge `https://www.anchorterminal.com/badges/granite-guardian.svg` or a link to https://www.anchorterminal.com/tools/granite-guardian from a page on ibm.com or one of its subdomains, or the README of github.com/ibm-granite/granite-guardian, then `POST https://www.anchorterminal.com/api/v1/verify` `{"slug", "url"}` or `verify_listing` at /mcp; re-checked weekly, no effect on the grade. Snippets in the full twin. ## Before you call it 1. Append the `` block as the final user message, with the mode line, `### Criteria:` and `### Scoring Schema:`. Copy the strings from the model card, because no package builds them 2. Use the no-think instruction for gating and parse ``. Think mode writes a reasoning trace first, and the card's examples allow up to 2,048 output tokens 3. Treat `yes` as the criterion being met, which for built-in criteria means the risk is present. Treat a missing `` tag as a failed check 4. Pass retrieved text through `documents=` and tool schemas through `available_tools=` in `apply_chat_template`, not inside the message text 5. Under Ollama, set `num_ctx` in the request options. IBM's docs say the default context is short and long requests are truncated ## Connect ```bash pip install transformers torch vllm ``` ```bash curl http://localhost:11434/api/chat \ -d '{"model": "granite4.1-guardian:8b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` ## Similar tools | Tool | Grade | Score | Shared capabilities | Slim | | --- | --- | --- | --- | --- | | OpenAI Guardrails | B | 69.5 | guard.injection, guard.moderation, guard.policy, guard.self-host | https://www.anchorterminal.com/tools/openai-guardrails.min.md | | NVIDIA NeMo Guardrails | B | 68.4 | guard.injection, guard.moderation, guard.policy, guard.self-host | https://www.anchorterminal.com/tools/nemo-guardrails.min.md | | Lakera Guard (Check Point AI Guardrails) | C | 59.6 | guard.injection, guard.moderation, guard.policy, guard.self-host | https://www.anchorterminal.com/tools/lakera-guard.min.md | | Guardrails AI | D | 49.6 | guard.injection, guard.moderation, guard.policy, guard.self-host | https://www.anchorterminal.com/tools/guardrails-ai.min.md | | Google Cloud Model Armor | BB | 77.9 | guard.injection, guard.moderation, guard.policy | https://www.anchorterminal.com/tools/google-model-armor.min.md | ## Panel reviews (0, desk reviews from public material, no calls made)