Head to head · Guard injection · October 2026 research run

Granite Guardian vs LlamaFirewall

Granite Guardian scores 60.1 (C) on agent readiness against LlamaFirewall's 50.8 (D), and leads in every scored category. Both do guard injection.

Which one, for what

Granite Guardian C

Good for A team with a GPU that wants one English-language judge for harm, jailbreaks, RAG groundedness, function-call checks and house rules, under a permissive licence with no gate.

Ahead on

  • Schema & documentation, 60 against 49
  • Agent ergonomics, 69 against 60
  • Payments & pricing, 60 against 50
  • Maintenance & community, 47 against 15
  • Transparency & trust, 70 against 58

Watch for

Trained and tested on English only, per the model card

LlamaFirewall D

Good for A Python agent team that wants injection, hidden-character and generated-code checks in process, is willing to pin dependencies or install from main, and can get the gated weights.

Also in its favour

  • Open source

Watch for

No PyPI release since 1.0.3 on 29 May 2025, and no changelog, tags or deprecation notes were found

Score by category

CategoryWeight this runGranite GuardianLlamaFirewallEdge
Reliability16%205653Granite Guardian +3
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.26049Granite Guardian +11
Agent ergonomics13%16.26960Granite Guardian +9
Security & auth14%17.55856Granite Guardian +2
Payments & pricing10%12.56050Granite Guardian +10
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.84715Granite Guardian +32
Transparency & trust7%8.87058Granite Guardian +12
Negative events≤1500
Total60.1 · C50.8 · D

Facts side by side

FactGranite GuardianLlamaFirewall
KindModel APIAgent framework
VendorIBMMeta
Hosted endpointno (local only)no (local only)
TransportsHTTP
AuthNoneNone
PricingFreeFree
x402nono
LicenceApache 2.0 for the weights and the repositoryMIT (library). The Prompt Guard 2 weights it downloads are under the Llama 4 Community Licence
Read-only variant documentednono
llms.txtnono
Last release2026-04-292025-05-29
Terms last updatedno document linkedno document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train models
Terms restrict automated access
Terms restrict benchmarking
Terms or service can change without notice
Arbitration or class-action waiver
Popularity182 stars4.4k stars, 1k PyPI/wk

Verdicts

Granite Guardian

One ungated Apache 2.0 model judges harm, jailbreaks, RAG groundedness, function-call errors and custom criteria, with signed weights and published evaluation code. It is trained and tested on English only, each call checks one criterion, the 4.1 prompt format differs from 3.x, and IBM's watsonx.ai lists only the deprecated 3.0 model.

LlamaFirewall

One scan() call runs several checks on the owner's machine and returns a short typed result. The last PyPI release is 1.0.3 from 29 May 2025, and its Prompt Guard loader imports a huggingface_hub class that current versions no longer export, so a fresh install needs older pins. The classifier weights also need Meta's manual approval.

Before you call either

Granite Guardian

  1. Append the <guardian> block as the final user message, with the mode line, ### Criteria: and ### Scoring Schema:. Copy the strings from the model card, because no package builds them
  2. Use the no-think instruction for gating and parse <score>. Think mode writes a reasoning trace first, and the card's examples allow up to 2,048 output tokens
  3. Treat yes as the criterion being met, which for built-in criteria means the risk is present. Treat a missing <score> tag as a failed check
  4. Pass retrieved text through documents= and tool schemas through available_tools= in apply_chat_template, not inside the message text
  5. Under Ollama, set num_ctx in the request options. IBM's docs say the default context is short and long requests are truncated

LlamaFirewall

  1. Pin huggingface_hub below 1.0 and a matching transformers 4.x before importing the Prompt Guard scanner from the 1.0.3 wheel, or install from main
  2. Get access to meta-llama/Llama-Prompt-Guard-2-86M and set a Hugging Face token first. Without one the loader prompts for a login and a headless run stalls
  3. Call scan_async inside a running event loop. scan() wraps asyncio.run and fails there. scan_async returns score 0.0 and reason default on every allow
  4. Split text longer than 512 tokens yourself before a Prompt Guard scan. The library truncates and does not chunk
  5. Do not feed a block reason back to the model. The Prompt Guard reason quotes the full scanned text, and the hidden ASCII reason decodes the hidden payload

Questions

Which is better for AI agents, Granite Guardian or LlamaFirewall?

Granite Guardian scores 60.1 (C) on agent readiness against LlamaFirewall's 50.8 (D), and leads in every scored category.

Can an agent call Granite Guardian and LlamaFirewall without installing anything?

No hosted endpoint is listed for Granite Guardian. No hosted endpoint is listed for LlamaFirewall.

Are Granite Guardian and LlamaFirewall open source?

No open-source release is listed for Granite Guardian. LlamaFirewall is open source (MIT (library). The Prompt Guard 2 weights it downloads are under the Llama 4 Community Licence).

Other comparisons with Granite Guardian or LlamaFirewall

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.