Head to head · Guard policy · October 2026 research run

Llama Guard 4 vs LlamaFirewall

LlamaFirewall scores 50.8 (D) on agent readiness against Llama Guard 4's 49.1 (D), and leads in 4 of 7 scored categories. Llama Guard 4 leads on agent ergonomics and maintenance & community. Both do guard policy.

Which one, for what

Llama Guard 4 D

Good for A team outside the EU with a GPU that wants content moderation of text and images on its own hardware, against a fixed 14-category policy it can edit in the prompt.

Ahead on

  • Agent ergonomics, 67 against 60
  • Maintenance & community, 28 against 15

Watch for

Weights last changed on 29 April 2025, with no changelog, version tags or stated deprecation policy

LlamaFirewall D

Good for A Python agent team that wants injection, hidden-character and generated-code checks in process, is willing to pin dependencies or install from main, and can get the gated weights.

Ahead on

  • Reliability, 53 against 38
  • Payments & pricing, 50 against 45

Also in its favour

  • Open source

Watch for

No PyPI release since 1.0.3 on 29 May 2025, and no changelog, tags or deprecation notes were found

Score by category

CategoryWeight this runLlama Guard 4LlamaFirewallEdge
Reliability16%203853LlamaFirewall +15
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25249Llama Guard 4 +3
Agent ergonomics13%16.26760Llama Guard 4 +7
Security & auth14%17.55356LlamaFirewall +3
Payments & pricing10%12.54550LlamaFirewall +5
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.82815Llama Guard 4 +13
Transparency & trust7%8.85558LlamaFirewall +3
Negative events≤1500
Total49.1 · D50.8 · D

Facts side by side

FactLlama Guard 4LlamaFirewall
KindModel APIAgent framework
VendorMetaMeta
Hosted endpointno (local only)no (local only)
TransportsHTTP
AuthNoneNone
PricingFreeFree
x402nono
LicenceLlama 4 Community Licence (source-available weights, not an OSI licence), with the Llama 4 acceptable use policyMIT (library). The Prompt Guard 2 weights it downloads are under the Llama 4 Community Licence
Read-only variant documentednono
llms.txtnono
Last release2025-04-292025-05-29
Terms last updated2025-04-05no document linked
Privacy policy last updatedno document linkedno document linked
Customer content may train modelsnot found in the text
Terms restrict automated accessnot found in the text
Terms restrict benchmarkingnot found in the text
Terms or service can change without noticenot found in the text
Arbitration or class-action waivernot found in the text
Popularity4.4k stars4.4k stars, 1k PyPI/wk

Verdicts

Llama Guard 4

A single self-hosted model classifies text and multi-image prompts against 14 MLCommons-aligned hazard categories and answers in a few tokens. The weights have not changed since 29 April 2025, download access needs Meta's manual approval, and the licence withholds the grant from individuals and companies based in the European Union.

LlamaFirewall

One scan() call runs several checks on the owner's machine and returns a short typed result. The last PyPI release is 1.0.3 from 29 May 2025, and its Prompt Guard loader imports a huggingface_hub class that current versions no longer export, so a fresh install needs older pins. The classifier weights also need Meta's manual approval.

Before you call either

Llama Guard 4

  1. Request access on the Hugging Face page before anything else. Approval is manual, and the form cannot be edited after submission
  2. Send only the user turn to check an input, and the user turn plus the model's answer to check an output. The template picks the role from the message count
  3. Parse the first line for safe or unsafe and the second for category codes. Set max_new_tokens to about 10 and turn sampling off
  4. Do not send an image with no text. Meta says the model is not an image-only classifier, and S14 is skipped when an image is present
  5. Pair it with a prompt-attack detector. The card says the model can itself be moved by adversarial or injected text

LlamaFirewall

  1. Pin huggingface_hub below 1.0 and a matching transformers 4.x before importing the Prompt Guard scanner from the 1.0.3 wheel, or install from main
  2. Get access to meta-llama/Llama-Prompt-Guard-2-86M and set a Hugging Face token first. Without one the loader prompts for a login and a headless run stalls
  3. Call scan_async inside a running event loop. scan() wraps asyncio.run and fails there. scan_async returns score 0.0 and reason default on every allow
  4. Split text longer than 512 tokens yourself before a Prompt Guard scan. The library truncates and does not chunk
  5. Do not feed a block reason back to the model. The Prompt Guard reason quotes the full scanned text, and the hidden ASCII reason decodes the hidden payload

Questions

Which is better for AI agents, Llama Guard 4 or LlamaFirewall?

LlamaFirewall scores 50.8 (D) on agent readiness against Llama Guard 4's 49.1 (D), and leads in 4 of 7 scored categories. Llama Guard 4 leads on agent ergonomics and maintenance & community.

Can an agent call Llama Guard 4 and LlamaFirewall without installing anything?

No hosted endpoint is listed for Llama Guard 4. No hosted endpoint is listed for LlamaFirewall.

Are Llama Guard 4 and LlamaFirewall open source?

No open-source release is listed for Llama Guard 4. LlamaFirewall is open source (MIT (library). The Prompt Guard 2 weights it downloads are under the Llama 4 Community Licence).

Other comparisons with Llama Guard 4 or LlamaFirewall

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.