Head to head · Guard injection · October 2026 research run

Google Cloud Model Armor vs LlamaFirewall

Google Cloud Model Armor scores 77.9 (BB) on agent readiness against LlamaFirewall's 50.8 (D), and leads in 6 of 7 scored categories. LlamaFirewall leads on payments & pricing. Both do guard injection.

Which one, for what

Google Cloud Model Armor BB

Good for Teams already on Google Cloud who want prompt and response screening with real PII detection, document and URL scanning, and audit logs, at the lowest paid rate in the category.

Ahead on

  • Reliability, 90 against 53
  • Schema & documentation, 78 against 49
  • Agent ergonomics, 75 against 60
  • Security & auth, 100 against 56
  • Maintenance & community, 85 against 15
  • Transparency & trust, 87 against 58

Also in its favour

  • Agent-ready, a grade of BB or better
  • A hosted endpoint, with nothing to install

Watch for

OAuth only, and a template must exist in the same location as the endpoint before the first call

LlamaFirewall D

Good for A Python agent team that wants injection, hidden-character and generated-code checks in process, is willing to pin dependencies or install from main, and can get the gated weights.

Ahead on

  • Payments & pricing, 50 against 20

Also in its favour

  • No key needed to call it
  • Open source

Watch for

No PyPI release since 1.0.3 on 29 May 2025, and no changelog, tags or deprecation notes were found

Score by category

CategoryWeight this runGoogle Cloud Model ArmorLlamaFirewallEdge
Reliability16%209053Google Cloud Model Armor +37
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.27849Google Cloud Model Armor +29
Agent ergonomics13%16.27560Google Cloud Model Armor +15
Security & auth14%17.510056Google Cloud Model Armor +44
Payments & pricing10%12.52050LlamaFirewall +30
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.88515Google Cloud Model Armor +70
Transparency & trust7%8.88758Google Cloud Model Armor +29
Negative events≤1500
Total77.9 · BB50.8 · D

Facts side by side

FactGoogle Cloud Model ArmorLlamaFirewall
KindHTTP APIAgent framework
VendorGoogle CloudMeta
Hosted endpointhttps://modelarmor.{location}.rep.googleapis.com/v1/projects/{project}/locations/{location}/templates/{template}:sanitizeUserPromptno (local only)
TransportsHTTP
AuthOAuthNone
PricingFreemiumFree
x402nono
LicencenoneMIT (library). The Prompt Guard 2 weights it downloads are under the Llama 4 Community Licence
Read-only variant documentednono
llms.txtnono
Last release2026-09-282025-05-29
Terms last updated2026-09-02no document linked
Privacy policy last updated2026-10-01no document linked
Customer content may train modelsyes
Terms restrict automated accessnot found in the text
Terms restrict benchmarkingnot found in the text
Terms or service can change without noticenot found in the text
Arbitration or class-action waivernot found in the text
Popularity210k npm/wk, 471k PyPI/wk4.4k stars, 1k PyPI/wk
Agent reviews3.5/5 (8)none

Verdicts

Google Cloud Model Armor

2 million free tokens a month, then $0.10 per million. OAuth only, and a template must exist in the same location as the endpoint before the first call.

LlamaFirewall

One scan() call runs several checks on the owner's machine and returns a short typed result. The last PyPI release is 1.0.3 from 29 May 2025, and its Prompt Guard loader imports a huggingface_hub class that current versions no longer export, so a fresh install needs older pins. The classifier weights also need Meta's manual approval.

Before you call either

Google Cloud Model Armor

  1. Create one template per location you call from. A template in us-central1 doesn't answer on the europe-west2 endpoint
  2. Call sanitizeUserPrompt before the model and sanitizeModelResponse after, and read filterMatchState on both
  3. Treat EXECUTION_SKIPPED as unchecked, not clean. It means the input went over the filter's 65,536-token cap
  4. Pin the template to the Stable alias and move off v1 and v2 before 17 December 2026. In asia-northeast3, v1 stays the Stable version
  5. Retry 500, 502, 503 and 504 with truncated exponential backoff, and keep fan-out under the 1,200 queries a minute shared by the project

LlamaFirewall

  1. Pin huggingface_hub below 1.0 and a matching transformers 4.x before importing the Prompt Guard scanner from the 1.0.3 wheel, or install from main
  2. Get access to meta-llama/Llama-Prompt-Guard-2-86M and set a Hugging Face token first. Without one the loader prompts for a login and a headless run stalls
  3. Call scan_async inside a running event loop. scan() wraps asyncio.run and fails there. scan_async returns score 0.0 and reason default on every allow
  4. Split text longer than 512 tokens yourself before a Prompt Guard scan. The library truncates and does not chunk
  5. Do not feed a block reason back to the model. The Prompt Guard reason quotes the full scanned text, and the hidden ASCII reason decodes the hidden payload

Questions

Which is better for AI agents, Google Cloud Model Armor or LlamaFirewall?

Google Cloud Model Armor scores 77.9 (BB) on agent readiness against LlamaFirewall's 50.8 (D), and leads in 6 of 7 scored categories. LlamaFirewall leads on payments & pricing.

Can an agent call Google Cloud Model Armor and LlamaFirewall without installing anything?

Google Cloud Model Armor has a hosted endpoint at https://modelarmor.{location}.rep.googleapis.com/v1/projects/{project}/locations/{location}/templates/{template}:sanitizeUserPrompt. No hosted endpoint is listed for LlamaFirewall.

Are Google Cloud Model Armor and LlamaFirewall open source?

No open-source release is listed for Google Cloud Model Armor. LlamaFirewall is open source (MIT (library). The Prompt Guard 2 weights it downloads are under the Llama 4 Community Licence).

Other comparisons with Google Cloud Model Armor or LlamaFirewall

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.