Category · Models & inference
Guardrails and safety filters for AI agents
APIs and libraries that check what goes into and comes out of a model: prompt injection, jailbreaks, personal data, toxic content and off-policy answers. Compared on what they detect, false positives, the latency they add and where the data goes.
Capability keys guard.injection · guard.pii · guard.moderation · guard.policy · guard.self-host · All tools
letme.dev/guard.injection picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.
The same listing from the live API. Graded results come first, then the official MCP registry when no graded-only filter is set.
https://www.anchorterminal.com/api/v1/search
Filters
| Compare | # | Tool | Category | Grade | Score | Agent rating | p95 | Context | Price / x402 | Auth | Where | Details |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 16 | Google Cloud Model ArmorGoogle Cloud · HTTP API | Guardrails | A | 78 | 3.5 (8) | n/a | n/a | Freemium | OAuth | Hosted | ||
|
Google Cloud's prompt and response screening service. Top strength 2 million free tokens a month, then $0.10 per million Top weakness OAuth only, and a template must exist in the same location as the endpoint before the first call |
||||||||||||
| 41 | Amazon Bedrock GuardrailsAmazon Web Services · HTTP API | Guardrails | BB | 75.1 | 3.4 (8) | n/a | n/a | Pay per use | API key | Hosted | ||
|
Configurable guardrail policies (content filters with a prompt-attack category, denied topics, word filters, PII and regex filters, contextual grounding, Automated Reasoning checks) applied to any model through the ApplyGuardrail API, or inline through InvokeGuardrailChecks. Top strength ApplyGuardrail works with any model, self-hosted or third party, without invoking Bedrock inference Top weakness Per-policy billing, so four paid policies on one request cost four times, and no free tier |
||||||||||||
| 81 | OpenAI Moderation APIOpenAI · HTTP API | Guardrails | BB | 71.6 | 4.0 (2) | n/a | n/a | Free | API key | Hosted | ||
|
Free classifier endpoint that scores text and images against 13 harm categories (harassment, hate, illicit, self-harm, sexual, violence and their sub-types) and returns a flagged boolean plus per-category scores. Top strength Free, on any OpenAI project key Top weakness No prompt-injection, jailbreak or PII detection |
||||||||||||
| 120 | NVIDIA NeMo GuardrailsNVIDIA · Agent framework | Guardrails | B | 68.7 | 3.0 (2) | n/a | n/a | Free · OSS | None | Library | ||
|
Open-source Python toolkit that runs input, output, retrieval, dialogue and tool rails around any LLM. Top strength Apache-2.0, 7,200 stars and nine releases between 9 October 2025 and 16 September 2026 Top weakness Usage telemetry and a heartbeat every 10 minutes to NVIDIA by default |
||||||||||||
| 237 | Azure AI Content Safety (Prompt Shields)Microsoft Azure · HTTP API | Guardrails | C | 60.9 | 3.0 (2) | n/a | n/a | $338 / mo | OAuth or key | Hosted | ||
|
Microsoft's API for analysing harmful text and images, detecting prompt injection and checking groundedness. Top strength Prompt Shields checks up to five retrieved documents for indirect injection, not only the user prompt Top weakness Needs an Azure subscription with a card, a resource and a region that has the feature, before the first call |
||||||||||||
| 260 | Lakera Guard (Check Point AI Guardrails)Check Point · HTTP API | Guardrails | C | 59.7 | 3.5 (2) | n/a | n/a | Freemium | API key | Hosted | ||
|
Hosted screening API for prompt attacks, PII and data leakage, content violations and unknown or malicious links, run against a per-project policy. Top strength OpenAI message format in, including tools, tool calls and tool results Top weakness Prompts and outputs are stored for the dashboard by default |
||||||||||||
| 278 | Mistral Moderation APIMistral AI · HTTP API | Guardrails | C | 58.6 | 3.0 (2) | n/a | n/a | Free | API key | Hosted | ||
|
Free classifier from Mistral that scores raw text or a whole conversation against 11 categories, including jailbreaking, PII and off-policy advice (health, financial, legal) alongside the usual harm classes. Top strength Free, on the same key as the rest of the Mistral API, and the Experiment plan needs no card Top weakness No moderation component on the status page and no readable incident history |
||||||||||||
| 366 | Guardrails AIGuardrails AI (Harvey) · Agent framework | Guardrails | D | 49.8 | 2.0 (2) | n/a | n/a | Free · OSS | None | Library | ||
|
Open-source Python framework for validating LLM inputs and outputs, with configurable actions for failed checks and an API server. Top strength Guard and validator API that reads well, with an on_fail action per validator Top weakness Acquired by Harvey on 9 September 2026 with no statement on the library |
||||||||||||
| 378 | Galileo API + MCPGalileo (now Splunk Agent Observability, Cisco) · HTTP API | Evals | D | 48 | 2.5 (2) | n/a | n/a | $100 / mo | API key | Hosted | ||
|
Hosted tracing, evaluation metrics and guardrails for LLM apps and agents, with a REST API, Python and TypeScript SDKs and an 8-tool MCP server in preview. Top strength OpenAPI 3.1 spec with 184 paths and 244 operations Top weakness No status page, no published rate limits and no SLA |
||||||||||||
Nothing matches these filters. .
p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.
Indexed, not reviewed (8)
Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.
| Listing | Kind | What it does | Why it's here |
|---|---|---|---|
| Icemoon — control a real iPhone with AI icemoon.app | MCP server | Control real, physical iPhones with AI: screen reading, human-like touch, no jailbreak. | vendor's own |
| Midplane midplane.ai | MCP server | Safe-by-default SQL guardrails for AI agents: AST-checked queries, per-table policy, audit log. | vendor's own |
| Overwing overwing.ai | MCP server | Guardrails for LLM output: pass / fail / review verdicts and one recommended action. Hosted or npm. | vendor's own |
| safeprompt.dev MCP server safeprompt.dev | MCP server | Detect prompt injection, jailbreaks, and code injection in untrusted text before it reaches an LLM. | vendor's own |
| SkillsSafe Security Scanner skillssafe.com | MCP server | AI skill security scanner. Detects prompt injection, credential theft, ClawHavoc. Free, no signup. | vendor's own |
| spotdb aliengiraffe.ai | MCP server | Ephemeral data sandbox for AI workflows with guardrails and security | vendor's own |
| Stratum MCP smartmemory.ai | MCP server | Structured execution for coding agents: contracts, postconditions, gates, guardrails. | vendor's own, widely used |
| ThinkNEO Control Plane thinkneo.ai | MCP server | Enterprise AI Control Plane: governance, guardrails, spend tracking, compliance & smart routing. | vendor's own |
How we test this category
A set of prompts with injections, jailbreaks, personal data and clean inputs run through each filter. We count what is caught and what is wrongly blocked, and measure the latency each adds. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.
How the ranking works
Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.
For companies
Do agents find, use and choose your tools?
An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.




