Best of · Models & inference

Best guardrails and safety filters for AI agents

The 10 highest-scoring of 16 guardrails and safety filters on the Anchor benchmark, with a pick for each need and where each one falls short. Scores come from public evidence, re-checked as vendors change.

  • 16 ranked
  • 3 agent-ready
  • 9 hosted endpoints
  • Updated 8 October 2026

Top three

Picks by need

Worked out from the scores, prices and facts, so they change when the research does.

Highest score overall

Google Cloud Model Armor BB

BB, 77.9/100 on the benchmark.

Also Amazon Bedrock Guardrails, BB, 74.8/100.

Schema & documentation

Amazon Bedrock Guardrails BB

92/100 on schema & documentation, against 78 for the overall leader.

Agent ergonomics

Amazon Bedrock Guardrails BB

93/100 on agent ergonomics, against 75 for the overall leader.

Maintenance & community

OpenAI Guardrails B

89/100 on maintenance & community, against 85 for the overall leader.

A hosted MCP endpoint

Galileo API + MCP D

remote MCP server, nothing to install.

Self-hosting under an open licence

OpenAI Guardrails B

self-hosted, MIT licence.

Also NVIDIA NeMo Guardrails, self-hosted, Apache-2 licence.

The shortlist

#ToolGradeBest forPriceWhere
1 Google Cloud Model Armor
Google Cloud
BB 77.9 Teams already on Google Cloud who want prompt and response screening with real PII detection, document and URL scanning, and audit logs, at the lowest paid rate in the category. Freemium hosted
2 Amazon Bedrock Guardrails
Amazon Web Services
BB 74.8 A team already on AWS that wants one versioned policy covering topics, PII masking, grounding and prompt attacks in front of any model. Pay per use hosted
3 OpenAI Moderation API
OpenAI
BB 71.3 A free harm-category filter for an agent already on OpenAI. Free hosted
4 OpenAI Guardrails
OpenAI
B 69.5 Teams already on the OpenAI client or Agents SDK that want several checks from one config file with little code. Free · OSS library
5 NVIDIA NeMo Guardrails
NVIDIA
B 68.4 Teams that want to compose several checks (their own, NVIDIA's and third-party APIs) behind one OpenAI-compatible endpoint. Free · OSS library
6 Presidio
Data Privacy Stack
B 66 Detecting and masking personal data in prompts, outputs, logs and images on the owner's own machines, with detection tuned by entity, threshold and custom recognisers. Free · OSS local
7 Prisma AIRS AI Runtime Security API
Palo Alto Networks, Inc.
B 62.8 A company already buying Palo Alto Networks through Strata Cloud Manager that wants prompt, response and MCP tool scanning with DLP and URL filtering from the same vendor. Paid hosted
8 Cisco AI Defense Inspection API
Cisco Systems, Inc.
C 61.3 A company already buying Cisco security through Security Cloud Control that wants its own application to decide what to do with each verdict. Paid hosted
9 Azure AI Content Safety (Prompt Shields)
Microsoft Azure
C 60.7 An agent on Azure that retrieves documents and needs indirect-injection checks next to harm-category moderation. $338 / mo hosted
10 Granite Guardian
IBM
C 60.1 A team with a GPU that wants one English-language judge for harm, jailbreaks, RAG groundedness, function-call checks and house rules, under a permissive licence with no gate. Free local

6 more are ranked in the full table.

How to choose

  1. Injection and jailbreak coverageCheck which attack types are detected, including indirect injection in tool results, because agents read untrusted pages and documents that can carry instructions.
  2. Wrongful blocks on clean inputCheck the false positive rate on clean inputs, because a filter that blocks legitimate tool calls stops the agent's task without any attack taking place.
  3. Latency added to each callCheck the latency each check adds at your payload size, since guardrails run on every model call and the delay compounds across a multi-step agent run.
  4. Self-hosting optionCheck whether the filter can run in your own infrastructure, because a hosted check sends every prompt, including personal data, to another party.

How the benchmark tests this category. A set of prompts with injections, jailbreaks, personal data and clean inputs run through each filter. We count what is caught and what is wrongly blocked, and measure the latency each adds.

Each one in detail

#1

Google Cloud Model Armor

BB 77.9/100

Google Cloud's prompt and response screening service.

Verdict 2 million free tokens a month, then $0.10 per million. OAuth only, and a template must exist in the same location as the endpoint before the first call.

Choose it for Teams already on Google Cloud who want prompt and response screening with real PII detection, document and URL scanning, and audit logs, at the lowest paid rate in the category.

Strengths

  • 2 million free tokens a month, then $0.10 per million
  • No incidents for Model Armor on the Google Cloud status page in the last 90 days
  • Each screening method has its own IAM permission and writes Data Access audit logs once the operator enables them

Weaknesses

  • OAuth only, and a template must exist in the same location as the endpoint before the first call
  • Filter versions v1 and v2 retire on 17 December 2026, a date that moved from 29 November within the same month
  • No SLA listed for Model Armor

Price FreemiumAuth OAuthx402 nohosted

Full assessment

#2

Amazon Bedrock Guardrails

BB 74.8/100

Configurable guardrail policies (content filters with a prompt-attack category, denied topics, word filters, PII and regex filters, contextual grounding, Automated Reasoning checks) applied to any model through the ApplyGuardrail API, or inline through InvokeGuardrailChecks.

Verdict ApplyGuardrail works with any model, self-hosted or third party, without invoking Bedrock inference. Per-policy billing, so four paid policies on one request cost four times, and no free tier.

Choose it for A team already on AWS that wants one versioned policy covering topics, PII masking, grounding and prompt attacks in front of any model.

Strengths

  • ApplyGuardrail works with any model, self-hosted or third party, without invoking Bedrock inference
  • InvokeGuardrailChecks takes the checks inline and returns severity and confidence scores, so no guardrail resource is needed
  • IAM can grant bedrock:ApplyGuardrail on one guardrail ARN and nothing else, and calls land in CloudTrail as data events

Weaknesses

  • Per-policy billing, so four paid policies on one request cost four times, and no free tier
  • Classic tier covers English, French and Spanish only, and Standard tier uses cross-Region inference that can move prompts within a geography
  • Quota numbers are mostly in the Service Quotas console, with public figures only for two US regions

Price Pay per useAuth API keyx402 nohosted

Full assessment · Against #1, Google Cloud Model Armor

#3

OpenAI Moderation API

BB 71.3/100

Free classifier endpoint that scores text and images against 13 harm categories (harassment, hate, illicit, self-harm, sexual, violence and their sub-types) and returns a flagged boolean plus per-category scores.

Verdict Free, on any OpenAI project key. No prompt-injection, jailbreak or PII detection.

Choose it for A free harm-category filter for an agent already on OpenAI.

Strengths

  • Free, on any OpenAI project key
  • A restricted key can be limited to the moderation endpoint
  • Text and images in the same request, with per-category scores

Weaknesses

  • No prompt-injection, jailbreak or PII detection
  • One model snapshot from 26 September 2024, and scores can shift when the latest alias moves
  • Fixed categories with no custom policies or per-request category choice

Price FreeAuth API keyx402 nohosted

Full assessment · Against #1, Google Cloud Model Armor

#4

OpenAI Guardrails

B 69.5/100

OpenAI's open-source Python library that wraps the OpenAI client and runs configured checks on inputs, outputs and tool calls, including moderation, jailbreak, prompt injection, personal data, URL and off-topic checks. It is labelled a preview.

Verdict MIT-licensed wrapper that adds twelve configurable checks to OpenAI client calls from one JSON file, with tool-level injection checks for the Agents SDK. The README labels it a preview at version 0.3.3, and by default a check that fails to run is reported as passed unless raise_guardrail_errors=True is set.

Choose it for Teams already on the OpenAI client or Agents SDK that want several checks from one config file with little code.

Strengths

  • MIT licence, source on GitHub, and three PyPI releases in the 90 days to 8 October 2026 (0.3.0, 0.3.2, 0.3.3)
  • Twelve built-in checks set in one versioned JSON file across pre-flight, input and output stages
  • GuardrailAgent runs the prompt injection check before and after every tool call in the OpenAI Agents SDK

Weaknesses

  • By default a check that fails to run returns tripwire_triggered=False, so the request continues. Strict mode is opt-in
  • The README titles the package a preview, the version is 0.3.3, and no release was published between 15 December 2025 and 21 July 2026
  • With stream=True the output checks run alongside the stream, and the docs say violating content may appear briefly

Price Free · OSSAuth API keyx402 nolibrary

Full assessment · Against #1, Google Cloud Model Armor

#5

NVIDIA NeMo Guardrails

B 68.4/100

Open-source Python toolkit that runs input, output, retrieval, dialogue and tool rails around any LLM.

Verdict Apache-2.0, 7,200 stars and nine releases between 9 October 2025 and 16 September 2026. Usage telemetry and a heartbeat every 10 minutes to NVIDIA by default.

Choose it for Teams that want to compose several checks (their own, NVIDIA's and third-party APIs) behind one OpenAI-compatible endpoint.

Strengths

  • Apache-2.0, 7,200 stars and nine releases between 9 October 2025 and 16 September 2026
  • Input, output, retrieval, dialogue, tool-input and tool-output rails in one config
  • Adapters for about 20 hosted guardrail services plus NVIDIA's NemoGuard models

Weaknesses

  • Usage telemetry and a heartbeat every 10 minutes to NVIDIA by default
  • Six breaking changes in 0.24.0, and the project is still pre-1.0
  • No authentication on the server, by design

Price Free · OSSAuth Nonex402 nolibrary

Full assessment · Against #1, Google Cloud Model Armor

#6

Presidio

B 66/100

Open-source Python library and Docker services that detect personal data in text and images and replace, mask, hash or encrypt it. Created at Microsoft and run since June 2026 by the community organisation Data Privacy Stack.

Verdict MIT-licensed personal data detector with a public OpenAPI document, tests on Python 3.10 to 3.14 and about 1.2 million weekly PyPI downloads. The REST containers have no authentication, the project states no SLA or support, and it covers personal data only, with no prompt injection or content moderation checks.

Choose it for Detecting and masking personal data in prompts, outputs, logs and images on the owner's own machines, with detection tuned by entity, threshold and custom recognisers.

Strengths

  • MIT licence, source on GitHub, and nothing to buy. No account, key or card is needed to install or run it
  • OpenAPI 3.0 document for the analyser and anonymiser REST services, with request examples and 400 and 422 error shapes
  • CI runs each package on Python 3.10, 3.11, 3.12, 3.13 and 3.14, with CodeQL and Dependabot configured

Weaknesses

  • The REST containers have no authentication by design. The FAQ says to put a gateway or proxy in front
  • SUPPORT.md states no SLA and no official support. The project is run by volunteers since leaving Microsoft
  • One release in the 90 days to 8 October 2026 (2.2.364 on 22 July), and CHANGELOG.md has no section for it

Price Free · OSSAuth Nonex402 nolocal

Full assessment · Against #1, Google Cloud Model Armor

#7

Prisma AIRS AI Runtime Security API

B 62.8/100

Hosted scan API from Palo Alto Networks that checks prompts, model responses and tool calls for prompt injection, sensitive data, toxic content, malicious URLs and code, and off-topic content against a security profile. Also reachable as a remote MCP server.

Verdict A public OpenAPI document, ten detection types in one call, tool-call scanning, a remote MCP server, rotating API keys and OAuth roles suit teams already on Strata Cloud Manager. Access needs Software NGFW credits bought through sales, with no public price, trial or self-serve signup, and payloads flagged malicious are kept for up to 10 years.

Choose it for A company already buying Palo Alto Networks through Strata Cloud Manager that wants prompt, response and MCP tool scanning with DLP and URL filtering from the same vendor.

Strengths

  • One call scans a prompt, a response and an MCP tool event for up to ten detection types set by the security profile
  • Public OpenAPI 3.0.3 documents for the scan API (4 operations) and the management API (21 operations)
  • API keys carry a rotation period, expiry and revocation, and OAuth service accounts take custom roles with per-entity permissions

Weaknesses

  • No public price, free tier or trial. Capacity is bought as Software NGFW credits, in steps of 1 billion tokens a month, through sales
  • Payloads judged malicious are kept for up to 10 years, including after the subscription ends, per the privacy datasheet of 13 July 2026
  • The EULA of August 2026 forbids publishing benchmark or comparison tests and forbids load testing of subscriptions

Price PaidAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Google Cloud Model Armor

#8

Cisco AI Defense Inspection API

C 61.3/100

Hosted inspection API from Cisco that checks chat messages, HTTP requests and responses, and MCP messages for prompt injection, personal data, harmful content and policy violations, and returns an allow or block verdict for the calling application to enforce.

Verdict Per-connection API keys with expiry, revocation and regeneration, published limits of 60,000 calls a minute, weekly dated release notes and a one-year retention statement suit companies already on Cisco Security Cloud Control. Access needs a subscription bought through sales, with no public price or trial for the API, and the developer changelog has one entry from February 2025.

Choose it for A company already buying Cisco security through Security Cloud Control that wants its own application to decide what to do with each verdict.

Strengths

  • One call returns is_safe, an action, classifications, severity, matched rules and, since April 2026, the type and position of each suspected PII value
  • Each connection has its own API key with an expiry date, revocation and regeneration, and the key works only on the Inspection API
  • Published limits of 60,000 calls a minute, a one million token context per inspection and one million tokens in parallel per organisation

Weaknesses

  • No public price, free tier or trial for the Inspection API. Subscriptions are bought through sales and activated with a claim code
  • The Toxicity rule was removed on 16 September 2026 in the release note that announced it, with no earlier notice found
  • The developer changelog lists only v1.0.0 of 28 February 2025, while release notes record later changes to the API's responses and rules

Price PaidAuth API keyx402 nohosted

Full assessment · Against #1, Google Cloud Model Armor

#9

Azure AI Content Safety (Prompt Shields)

C 60.7/100

Microsoft's API for analysing harmful text and images, detecting prompt injection and checking groundedness.

Verdict Prompt Shields checks up to five retrieved documents for indirect injection, not only the user prompt. Needs an Azure subscription with a card, a resource and a region that has the feature, before the first call.

Choose it for An agent on Azure that retrieves documents and needs indirect-injection checks next to harm-category moderation.

Strengths

  • Prompt Shields checks up to five retrieved documents for indirect injection, not only the user prompt
  • 5,000 free text records and 5,000 free images a month on F0
  • FAQ and data-privacy page agree that inputs aren't stored or trained on and stay in the resource's region

Weaknesses

  • Needs an Azure subscription with a card, a resource and a region that has the feature, before the first call
  • Python SDK is 1.0.0 from December 2023 and has no Prompt Shields method
  • No retry or 429 guidance in the Content Safety docs

Price $338 / moAuth OAuth or keyx402 nohosted

Full assessment · Against #1, Google Cloud Model Armor

#10

Granite Guardian

C 60.1/100

Granite Guardian is IBM's family of open-weight judge models. The current 8-billion-parameter release answers yes or no on whether a prompt, response, retrieved context or function call meets a built-in or custom criterion, and the owner runs it.

Verdict One ungated Apache 2.0 model judges harm, jailbreaks, RAG groundedness, function-call errors and custom criteria, with signed weights and published evaluation code. It is trained and tested on English only, each call checks one criterion, the 4.1 prompt format differs from 3.x, and IBM's watsonx.ai lists only the deprecated 3.0 model.

Choose it for A team with a GPU that wants one English-language judge for harm, jailbreaks, RAG groundedness, function-call checks and house rules, under a permissive licence with no gate.

Strengths

  • Weights are ungated on Hugging Face under the Apache 2.0 licence, with IBM's own GGUF builds and an Ollama library entry
  • Built-in criteria cover harm, social bias, jailbreaking, violence, profanity, sexual content, unethical behaviour, three RAG checks and function-call hallucination
  • A custom criterion is one natural-language sentence in the prompt, and the answer is yes or no inside <score> tags

Weaknesses

  • Trained and tested on English only, per the model card
  • Each call judges one criterion, so checking several risks takes several calls or a separate LoRA adapter built on the 3.2 model
  • Version 4.1 moved the criterion into a <guardian> block in the last user message, where 3.x cookbooks pass guardian_config

Price FreeAuth Nonex402 nolocal

Full assessment · Against #1, Google Cloud Model Armor

Head to head

All 118 comparisons in this category

Questions

What are the highest-rated guardrails and safety filters for AI agents?

Google Cloud Model Armor has the highest benchmark score of the 16 ranked guardrails and safety filters, 77.9 (BB). Amazon Bedrock Guardrails is second with 74.8 (BB).

How many guardrails and safety filters are agent-ready?

3 of the 16 ranked here grade BB or better, the bar for agent-ready on the Anchor benchmark.

Which guardrails and safety filters accept x402 payments?

None of the ranked listings here accepts x402 for its main call yet.

How is this list ranked?

By the Anchor benchmark score out of 100, a weighted mean of the scored categories minus deductions for negative events, from public evidence re-checked as vendors change. Listings cannot pay for a place. The latest assessment behind this page is from 8 October 2026.

How this list is made

The order is the Anchor benchmark score, the same number as on each listing and in the top list. Each listing is graded from public evidence against the benchmark checklist, and the picks above are worked out from those grades, prices and facts. No listing pays for its place, and paid audits or listing help never change a score.

Full ranked table · 118 head-to-head comparisons · Best tools in every category

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.