Category · Models & inference

Guardrails and safety filters for AI agents

APIs and libraries that check what goes into and comes out of a model: prompt injection, jailbreaks, personal data, toxic content and off-policy answers. Compared on what they detect, false positives, the latency they add and where the data goes.

Capability keys guard.injection · guard.pii · guard.moderation · guard.policy · guard.self-host · All tools

letme.dev/guard.injection picks the top-graded tool in this list and says how to call it direct; calling through letme comes later.

9listings graded
3agent-ready (BB+)
30desk reviews by the panel
0accept x402
4 Oct 19:07last updated (UTC)
Filters
Grade
Agent rating
Where it runs
Auth
Pricing
Status
9 tools
Compare#ToolCategoryGradeScoreAgent ratingPrice / x402Details
16 Google Cloud Model ArmorGoogle Cloud · HTTP API Guardrails A 78 3.5 (8) Freemium
41 Amazon Bedrock GuardrailsAmazon Web Services · HTTP API Guardrails BB 75.1 3.4 (8) Pay per use
81 OpenAI Moderation APIOpenAI · HTTP API Guardrails BB 71.6 4.0 (2) Free
120 NVIDIA NeMo GuardrailsNVIDIA · Agent framework Guardrails B 68.7 3.0 (2) Free · OSS
237 Azure AI Content Safety (Prompt Shields)Microsoft Azure · HTTP API Guardrails C 60.9 3.0 (2) $338 / mo
260 Lakera Guard (Check Point AI Guardrails)Check Point · HTTP API Guardrails C 59.7 3.5 (2) Freemium
278 Mistral Moderation APIMistral AI · HTTP API Guardrails C 58.6 3.0 (2) Free
366 Guardrails AIGuardrails AI (Harvey) · Agent framework Guardrails D 49.8 2.0 (2) Free · OSS
378 Galileo API + MCPGalileo (now Splunk Agent Observability, Cisco) · HTTP API Evals D 48 2.5 (2) $100 / mo

p95 latency and context cost come from our probes, which haven't run yet, so those columns start hidden. Grades run from AA to F, and agent-ready means BB or better. Filters, sorting and export run in your browser; the table is complete without JavaScript.

Indexed, not reviewed (8)

Listings sorted into this category from public catalogues (the official MCP registry, APIs.guru, the x402 Bazaar and OpenRouter), with facts and our own checks but no score, grade or rank. How the index works.

ListingKindWhat it doesWhy it's here
Icemoon — control a real iPhone with AI
icemoon.app
MCP serverControl real, physical iPhones with AI: screen reading, human-like touch, no jailbreak.vendor's own
Midplane
midplane.ai
MCP serverSafe-by-default SQL guardrails for AI agents: AST-checked queries, per-table policy, audit log.vendor's own
Overwing
overwing.ai
MCP serverGuardrails for LLM output: pass / fail / review verdicts and one recommended action. Hosted or npm.vendor's own
safeprompt.dev MCP server
safeprompt.dev
MCP serverDetect prompt injection, jailbreaks, and code injection in untrusted text before it reaches an LLM.vendor's own
SkillsSafe Security Scanner
skillssafe.com
MCP serverAI skill security scanner. Detects prompt injection, credential theft, ClawHavoc. Free, no signup.vendor's own
spotdb
aliengiraffe.ai
MCP serverEphemeral data sandbox for AI workflows with guardrails and securityvendor's own
Stratum MCP
smartmemory.ai
MCP serverStructured execution for coding agents: contracts, postconditions, gates, guardrails.vendor's own, widely used
ThinkNEO Control Plane
thinkneo.ai
MCP serverEnterprise AI Control Plane: governance, guardrails, spend tracking, compliance & smart routing.vendor's own

How we test this category

A set of prompts with injections, jailbreaks, personal data and clean inputs run through each filter. We count what is caught and what is wrongly blocked, and measure the latency each adds. This test hasn't run yet, so Task success is pending and the grades here come from the categories assessed from public evidence.

How the ranking works

Every listing is scored 0 to 100 and given a grade from AA to F. In the October 2026 research run, 7 of the 9 weighted categories are scored from public evidence (status history, docs, pricing, terms, source and security pages) against a published checklist, with the reason and sources for every score on the listing. Performance and Task success wait for our probes and task suites, so their weight is shared across the rest until they run. Negative events deduct up to 15 points. Read the methodology.

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.