# Guardrails and safety filters for AI agents > 9 guardrails & safety filters ranked by the Anchor benchmark. Leader Google Cloud Model Armor (A). APIs and libraries that check what goes into and comes out of a model: prompt injection, jailbreaks, personal data, toxic content and off-policy answers. Compared on what they detect, false positives, the latency they add and where the data goes. - Canonical: https://www.anchorterminal.com/categories/guardrails - Markdown: https://www.anchorterminal.com/categories/guardrails.md (~3,100 tokens) - Slim: https://www.anchorterminal.com/categories/guardrails.min.md (~530 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/categories/guardrails.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-04 APIs and libraries that check what goes into and comes out of a model: prompt injection, jailbreaks, personal data, toxic content and off-policy answers. Compared on what they detect, false positives, the latency they add and where the data goes. - Tools ranked: 9 · agent-ready (BB or better): 3 · accept x402: 0 · hosted endpoints: 7 · desk reviews by the panel: 30 - JSON: https://www.anchorterminal.com/api/v1/tools.json (list) · https://www.anchorterminal.com/api/v1/rankings.json (ranked) · https://www.anchorterminal.com/api/v1/x402.json (payable) · https://www.anchorterminal.com/api/v1/capabilities.json (by capability) - Grades run AA, A, BB, B, C, D, E, F · methodology: https://www.anchorterminal.com/benchmark/ - Capabilities in this category: guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host - https://letme.dev/guard.injection picks the top-graded tool in this list and says how to call it direct; calling through letme comes later (https://www.anchorterminal.com/letme/index.md) ## Ranking | # | Tool | Vendor | Kind | Category | Grade | Score | Confidence | x402 | Auth | Where | Reviews | Page | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 16 | Google Cloud Model Armor | Google Cloud | HTTP API | Guardrails | A | 78 | high | no | OAuth | hosted | 3.5/5 (8) | https://www.anchorterminal.com/tools/google-model-armor.md | | 41 | Amazon Bedrock Guardrails | Amazon Web Services | HTTP API | Guardrails | BB | 75.1 | medium | no | API key | hosted | 3.4/5 (8) | https://www.anchorterminal.com/tools/amazon-bedrock-guardrails.md | | 81 | OpenAI Moderation API | OpenAI | HTTP API | Guardrails | BB | 71.6 | high | no | API key | hosted | 4/5 (2) | https://www.anchorterminal.com/tools/openai-moderation.md | | 120 | NVIDIA NeMo Guardrails | NVIDIA | Agent framework | Guardrails | B | 68.7 | medium | no | None | library | 3/5 (2) | https://www.anchorterminal.com/tools/nemo-guardrails.md | | 237 | Azure AI Content Safety (Prompt Shields) | Microsoft Azure | HTTP API | Guardrails | C | 60.9 | medium | no | OAuth or key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/azure-ai-content-safety.md | | 260 | Lakera Guard (Check Point AI Guardrails) | Check Point | HTTP API | Guardrails | C | 59.7 | medium | no | API key | hosted | 3.5/5 (2) | https://www.anchorterminal.com/tools/lakera-guard.md | | 278 | Mistral Moderation API | Mistral AI | HTTP API | Guardrails | C | 58.6 | medium | no | API key | hosted | 3/5 (2) | https://www.anchorterminal.com/tools/mistral-moderation.md | | 366 | Guardrails AI | Guardrails AI (Harvey) | Agent framework | Guardrails | D | 49.8 | medium | no | None | library | 2/5 (2) | https://www.anchorterminal.com/tools/guardrails-ai.md | | 378 | Galileo API + MCP | Galileo (now Splunk Agent Observability, Cisco) | HTTP API | Evals | D | 48 | medium | no | API key | hosted | 2.5/5 (2) | https://www.anchorterminal.com/tools/galileo.md | Scores are from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/), with Performance and Task success pending. p95 latency and context cost come from our probes, which haven't run yet. ## Summaries ### 16. Google Cloud Model Armor, A (78) Google Cloud's prompt and response screening service. 2 million free tokens a month, then $0.10 per million. OAuth only, and a template must exist in the same location as the endpoint before the first call. - Page: https://www.anchorterminal.com/tools/google-model-armor · Markdown: https://www.anchorterminal.com/tools/google-model-armor.md · JSON: https://www.anchorterminal.com/api/v1/tools/google-model-armor.json - Capabilities: guard.injection, guard.pii, guard.moderation, guard.policy · endpoint: `https://modelarmor.{location}.rep.googleapis.com/v1/projects/{project}/locations/{location}/templates/{template}:sanitizeUserPrompt` ### 41. Amazon Bedrock Guardrails, BB (75.1) Configurable guardrail policies (content filters with a prompt-attack category, denied topics, word filters, PII and regex filters, contextual grounding, Automated Reasoning checks) applied to any model through the ApplyGuardrail API, or inline through InvokeGuardrailChecks. ApplyGuardrail works with any model, self-hosted or third party, without invoking Bedrock inference. Per-policy billing, so four paid policies on one request cost four times, and no free tier. - Page: https://www.anchorterminal.com/tools/amazon-bedrock-guardrails · Markdown: https://www.anchorterminal.com/tools/amazon-bedrock-guardrails.md · JSON: https://www.anchorterminal.com/api/v1/tools/amazon-bedrock-guardrails.json - Capabilities: guard.injection, guard.pii, guard.moderation, guard.policy · endpoint: `https://bedrock-runtime.{region}.amazonaws.com/guardrail/{id}/version/{version}/apply` ### 81. OpenAI Moderation API, BB (71.6) Free classifier endpoint that scores text and images against 13 harm categories (harassment, hate, illicit, self-harm, sexual, violence and their sub-types) and returns a flagged boolean plus per-category scores. Free, on any OpenAI project key. No prompt-injection, jailbreak or PII detection. - Page: https://www.anchorterminal.com/tools/openai-moderation · Markdown: https://www.anchorterminal.com/tools/openai-moderation.md · JSON: https://www.anchorterminal.com/api/v1/tools/openai-moderation.json - Capabilities: guard.moderation · endpoint: `https://api.openai.com/v1/moderations` ### 120. NVIDIA NeMo Guardrails, B (68.7) Open-source Python toolkit that runs input, output, retrieval, dialogue and tool rails around any LLM. Apache-2.0, 7,200 stars and nine releases between 9 October 2025 and 16 September 2026. Usage telemetry and a heartbeat every 10 minutes to NVIDIA by default. - Page: https://www.anchorterminal.com/tools/nemo-guardrails · Markdown: https://www.anchorterminal.com/tools/nemo-guardrails.md · JSON: https://www.anchorterminal.com/api/v1/tools/nemo-guardrails.json - Capabilities: guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host ### 237. Azure AI Content Safety (Prompt Shields), C (60.9) Microsoft's API for analysing harmful text and images, detecting prompt injection and checking groundedness. Prompt Shields checks up to five retrieved documents for indirect injection, not only the user prompt. Needs an Azure subscription with a card, a resource and a region that has the feature, before the first call. - Page: https://www.anchorterminal.com/tools/azure-ai-content-safety · Markdown: https://www.anchorterminal.com/tools/azure-ai-content-safety.md · JSON: https://www.anchorterminal.com/api/v1/tools/azure-ai-content-safety.json - Capabilities: guard.injection, guard.moderation, guard.policy · endpoint: `https://{resource}.cognitiveservices.azure.com/contentsafety/text:shieldPrompt` ### 260. Lakera Guard (Check Point AI Guardrails), C (59.7) Hosted screening API for prompt attacks, PII and data leakage, content violations and unknown or malicious links, run against a per-project policy. OpenAI message format in, including tools, tool calls and tool results. Prompts and outputs are stored for the dashboard by default. - Page: https://www.anchorterminal.com/tools/lakera-guard · Markdown: https://www.anchorterminal.com/tools/lakera-guard.md · JSON: https://www.anchorterminal.com/api/v1/tools/lakera-guard.json - Capabilities: guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host · endpoint: `https://api.lakera.ai/v2/guard` ### 278. Mistral Moderation API, C (58.6) Free classifier from Mistral that scores raw text or a whole conversation against 11 categories, including jailbreaking, PII and off-policy advice (health, financial, legal) alongside the usual harm classes. Free, on the same key as the rest of the Mistral API, and the Experiment plan needs no card. No moderation component on the status page and no readable incident history. - Page: https://www.anchorterminal.com/tools/mistral-moderation · Markdown: https://www.anchorterminal.com/tools/mistral-moderation.md · JSON: https://www.anchorterminal.com/api/v1/tools/mistral-moderation.json - Capabilities: guard.moderation, guard.pii, guard.policy · endpoint: `https://api.mistral.ai/v1/moderations` ### 366. Guardrails AI, D (49.8) Open-source Python framework for validating LLM inputs and outputs, with configurable actions for failed checks and an API server. Validators have configurable actions for failed checks. Harvey acquired the company on 9 September 2026; the reviewed announcement did not state plans for the library. - Page: https://www.anchorterminal.com/tools/guardrails-ai · Markdown: https://www.anchorterminal.com/tools/guardrails-ai.md · JSON: https://www.anchorterminal.com/api/v1/tools/guardrails-ai.json - Capabilities: guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host ### 378. Galileo API + MCP, D (48) Hosted tracing, evaluation metrics and guardrails for LLM apps and agents, with a REST API, Python and TypeScript SDKs and an 8-tool MCP server in preview. OpenAPI 3.1 spec with 184 paths and 244 operations. No status page, no published rate limits and no SLA. - Page: https://www.anchorterminal.com/tools/galileo · Markdown: https://www.anchorterminal.com/tools/galileo.md · JSON: https://www.anchorterminal.com/api/v1/tools/galileo.json - Capabilities: obs.traces, obs.evals, obs.prompts, obs.datasets · endpoint: `https://api.galileo.ai/v2` ## How we test this category A set of prompts with injections, jailbreaks, personal data and clean inputs run through each filter. We count what is caught and what is wrongly blocked, and measure the latency each adds. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence. ## Indexed, not reviewed (8) Sorted into this category from public catalogues, with facts and our own checks but no score, grade or rank (https://www.anchorterminal.com/indexed/index.md). | Listing | Kind | What it does | Why it's here | | --- | --- | --- | --- | | [Icemoon — control a real iPhone with AI](https://www.anchorterminal.com/tools/icemoon-automation-studio.md) | MCP server | Control real, physical iPhones with AI: screen reading, human-like touch, no jailbreak. | vendor's own | | [Midplane](https://www.anchorterminal.com/tools/midplane.md) | MCP server | Safe-by-default SQL guardrails for AI agents: AST-checked queries, per-table policy, audit log. | vendor's own | | [Overwing](https://www.anchorterminal.com/tools/overwing-mcp.md) | MCP server | Guardrails for LLM output: pass / fail / review verdicts and one recommended action. Hosted or npm. | vendor's own | | [safeprompt.dev MCP server](https://www.anchorterminal.com/tools/safeprompt-mcp.md) | MCP server | Detect prompt injection, jailbreaks, and code injection in untrusted text before it reaches an LLM. | vendor's own | | [SkillsSafe Security Scanner](https://www.anchorterminal.com/tools/skillssafe-scanner.md) | MCP server | AI skill security scanner. Detects prompt injection, credential theft, ClawHavoc. Free, no signup. | vendor's own | | [spotdb](https://www.anchorterminal.com/tools/aliengiraffe-spotdb.md) | MCP server | Ephemeral data sandbox for AI workflows with guardrails and security | vendor's own | | [Stratum MCP](https://www.anchorterminal.com/tools/smartmemory-stratum-mcp.md) | MCP server | Structured execution for coding agents: contracts, postconditions, gates, guardrails. | vendor's own, widely used | | [ThinkNEO Control Plane](https://www.anchorterminal.com/tools/thinkneo-control-plane.md) | MCP server | Enterprise AI Control Plane: governance, guardrails, spend tracking, compliance & smart routing. | vendor's own |