# OpenAI Guardrails > OpenAI's open-source Python library that wraps the OpenAI client and runs configured checks on inputs, outputs and tool calls, including moderation, jailbreak, prompt injection, personal data, URL and off-topic checks. It is labelled a preview. - Canonical: https://www.anchorterminal.com/tools/openai-guardrails - Markdown: https://www.anchorterminal.com/tools/openai-guardrails.md (~6,750 tokens) - Slim: https://www.anchorterminal.com/tools/openai-guardrails.min.md (~1,430 tokens, same facts, less prose, for token-sensitive contexts) - JSON: https://www.anchorterminal.com/tools/openai-guardrails.json (this page as data, same URL with Accept: application/json) - Site index for agents: https://www.anchorterminal.com/llms.txt (full text: https://www.anchorterminal.com/llms-full.txt) - API: https://www.anchorterminal.com/api/v1/index.json - Updated: 2026-10-09 ## Overview **Grade B · 69.5/100 · rank #175 of 842 · #4 in Guardrails & safety filters · not agent-ready · confidence medium** More from OpenAI, listed separately because each is its own product: [OpenAI API](https://www.anchorterminal.com/tools/openai-api.md) (Model APIs & inference), [OpenAI embeddings](https://www.anchorterminal.com/tools/openai-embeddings.md) (Embeddings & rerankers), [OpenAI Moderation API](https://www.anchorterminal.com/tools/openai-moderation.md) (Guardrails & safety filters), [OpenAI Image API](https://www.anchorterminal.com/tools/openai-image-api.md) (Image generation), [OpenAI Sora API](https://www.anchorterminal.com/tools/openai-sora.md) (Video generation), [OpenAI Agents SDK](https://www.anchorterminal.com/tools/openai-agents-sdk.md) (Agent frameworks & SDKs), [OpenAI Decisions API](https://www.anchorterminal.com/tools/openai-decisions-api.md) (Decision models), [OpenAI Codex](https://www.anchorterminal.com/tools/openai-codex.md) (Agent harnesses). ## Assessment MIT-licensed wrapper that adds twelve configurable checks to OpenAI client calls from one JSON file, with tool-level injection checks for the Agents SDK. The README labels it a preview at version 0.3.3, and by default a check that fails to run is reported as passed unless `raise_guardrail_errors=True` is set. ## Facts | Field | Value | | --- | --- | | Vendor | OpenAI (https://openai.com) | | Kind | Agent framework | | Category | Guardrails & safety filters (https://www.anchorterminal.com/categories/guardrails) | | Transport | HTTP | | Auth | API key · No credential of its own. The wrapped client takes an OpenAI API key from `OPENAI_API_KEY` or the constructor, an Azure OpenAI key through `GuardrailsAzureOpenAI`, or the key of any OpenAI-compatible endpoint set with `base_url`, such as a local Ollama server (https://openai.github.io/openai-guardrails-python/quickstart/). | | Pricing | Free (Free · OSS) · Free under the MIT licence, with no hosted or paid edition and no account of its own. The README states that Guardrails calls paid OpenAI APIs. Each LLM-based check is one extra model call billed at OpenAI's rates, the moderation check is documented as no cost, and the keyword, URL, secret key and personal data checks run locally (https://github.com/openai/openai-guardrails-python). | | x402 | No · No x402, MPP or L402 in the README or docs (checked 2026-10-08). | | Licence | MIT | | Packages | pypi: `openai-guardrails`; npm: `@openai/guardrails` | | Source | https://github.com/openai/openai-guardrails-python | | Docs | https://openai.github.io/openai-guardrails-python/ | | llms.txt | not found | | Last release | 2026-09-10 | | GitHub stars | 259 (as of 2026-10-08) | | npm downloads / week | 18,502 | | PyPI downloads / week | 104,530 | | Packages | `openai-guardrails` 0.3.3 on PyPI, `@openai/guardrails` 0.3.0 on npm | | Languages | Python 3.11 to 3.14. TypeScript on Node.js 22.13 or later | | Status | Preview, per the README title | | Clients | `GuardrailsOpenAI`, `GuardrailsAsyncOpenAI`, `GuardrailsAzureOpenAI`, `GuardrailsAsyncAzureOpenAI`, and `GuardrailAgent` for the OpenAI Agents SDK | | Wrapped calls | `chat.completions.create`, `responses.create` and `responses.parse` | | Stages | Pre-flight (before the model call), input (in parallel with it) and output | | Checks | Keyword Filter, Competitors, Moderation, URL Filter, Secret Keys, Contains PII, Hallucination Detection, Jailbreak, Prompt Injection Detection, NSFW Text, Off Topic Prompts, Custom Prompt Check | | Configuration | One versioned JSON file, a dict or a JSON string. A wizard at https://guardrails.openai.com/ exports the file | | On a violation | Raises `GuardrailTripwireTriggered`, or returns results on `response.guardrail_results` with `suppress_tripwire=True` | | On a check error | Passes by default. `raise_guardrail_errors=True` raises | | Command line | `guardrails validate ` and `guardrails-evals` for labelled datasets | | Telemetry | None found in the source (searched 2026-10-08). Calls to api.openai.com carry `safety_identifier` `openai-guardrails-python` | | Releases in 90 days | 3 (0.3.0 on 2026-07-21, 0.3.2 on 2026-08-21, 0.3.3 on 2026-09-10) | | Capabilities | guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host | | Tags | official, framework, open-source, self-hosted, local, python, typescript, free, preview, openai-compatible | | JSON | https://www.anchorterminal.com/api/v1/tools/openai-guardrails.json | ## Score breakdown (methodology v0.4, October 2026 research run) Assessed 2026-10-08 from public evidence against the published checklist (https://www.anchorterminal.com/benchmark/#checklist). Confidence: medium. Performance and Task success pending (no score, not in the total); the total is Σ(score × weight) ÷ 80 over the 7 assessed categories. "This run" is each category's share of the 100 points. | Category | Weight | This run | Score (0–100) | Points | | --- | --- | --- | --- | --- | | Reliability | 16% | 20 | 73 | 14.6 | | Performance | 10% | pending | pending | n/a | | Schema & documentation | 13% | 16.2 | 66 | 10.7 | | Agent ergonomics | 13% | 16.2 | 73 | 11.9 | | Security & auth | 14% | 17.5 | 59 | 10.3 | | Payments & pricing | 10% | 12.5 | 60 | 7.5 | | Task success | 10% | pending | pending | n/a | | Maintenance & community | 7% | 8.8 | 89 | 7.8 | | Transparency & trust (editorial 65, provenance 87) | 7% | 8.8 | 76 | 6.7 | | Negative events | up to −15 | up to −15 | none recorded | 0 | | **Total** | | | | **69.5 → B** | ### Why each score - Reliability 73: Local-software reading. Installs from PyPI as `openai-guardrails`, Python 3.11 to 3.14 stated (20). Public CI runs ruff, mypy, pyright and the test suite on four Python versions, and the CI run on main passed on 8 October 2026 (25). No open issues and 11 closed in total, though four were closed together on 18 August 2026 after three to ten months (20 of 25). Semver tags, but CHANGELOG.md starts at 0.3.3 and no release note marks a breaking change. RELEASING.md says breaking changes bump the minor version before 1.0 (8 of 15). Version 0.3.3 and the README title reads Preview (0). - Performance: Pending. Latency is measured per call by our probes, which haven't run yet, so this run doesn't score it. Its weight is shared across the assessed categories until the first probe window closes. - Schema & documentation 66: Framework reading. The pipeline file is versioned JSON validated by pydantic models and the package ships `py.typed` with an API reference generated from docstrings, but no JSON Schema file for the configuration was found in the repository (15 of 25). `llms.txt` on the docs site returns 404 (0). Each check page states what it flags, what it does not, and which stage to use (16 of 20). Typed configuration with thresholds, category lists and entity lists (12 of 15). Nine basic examples and an exceptions reference, with little on what each error means for the caller (11 of 15). Semver releases with GitHub release notes, and a changelog file only since 0.3.3 (12 of 15). The docs deploy workflow failed on main on 8 October 2026. - Agent ergonomics 73: Framework reading. `GuardrailsOpenAI` replaces `OpenAI` with one config argument, and results sit on `response.guardrail_results` (22 of 25). Confidence thresholds, `max_turns`, `include_reasoning` off by default and `suppress_tripwire` control what comes back, and `total_guardrail_token_usage` reports the cost (15 of 20). Typed exceptions carry the check name and stage, but a check that fails to run is reported as not triggered unless `raise_guardrail_errors=True` (12 of 20). Checks are stateless and safe to repeat, each LLM-based check is an extra model call, and streaming can show output before an output check trips (12 of 20). Python and TypeScript packages, few required parameters, and a spaCy model to install by hand for Contains PII (12 of 15). - Security & auth 59: Framework reading. No credential of its own. It uses the OpenAI API key, or the key of whichever compatible endpoint is set, from the environment or the constructor (10 of 30). The prompt injection check runs before and after each tool call in the Agents SDK and can reject or halt, but there is no human approval step, and check failures pass by default (10 of 20). Jailbreak and prompt injection detection are the product, with a published benchmark in which the default model scores 0.000 recall at a 1 per cent false positive rate (13 of 15). Per-call results with check name, stage, confidence and token usage, and no log sink of its own (9 of 15). SECURITY.md points to OpenAI's coordinated disclosure policy, security.txt names a Bugcrowd contact, CodeQL and Dependabot are configured, releases use PyPI trusted publishing with attestations, and the repository has no published advisories (17 of 20). - Payments & pricing 60: Self-hosted rule. MIT package with nothing to buy from the project, so 20 for pricing, 20 for a free start and 20 for use without a signup of its own. No payment protocol (0). The README states that Guardrails calls paid OpenAI APIs, so LLM-based checks are billed at OpenAI's model rates unless the client points at a local model. The moderation check page says that call has no cost. - Task success: Pending. Task success needs the category task suites run through each tool, which haven't run yet, so this run doesn't score it. Its weight is shared across the assessed categories until then. A data provider's data-quality score is published on its listing now and becomes half of this category when it's scored. - Maintenance & community 89: 0.3.3 on 10 September 2026, 28 days before the check (30). Three PyPI releases in the last 90 days, 0.3.0 on 21 July, 0.3.2 on 21 August and 0.3.3 (20). All 11 issues are closed and bug reports from 2025 were answered within days, but four were closed in bulk on 18 August 2026, outside pull requests are refused by policy, and nothing was released from 15 December 2025 to 21 July 2026 (15 of 25). Current official packages on PyPI and npm (15). Dependabot, CodeQL and CI on four Python versions, with the docs deploy failing on 8 October (9 of 10). - Transparency & trust 76: MIT licence in the repository and on PyPI (30). The README says the developer is responsible for storage of blocked content and that the library calls OpenAI APIs, and each check page says whether it uses a model, but no page lists which checks send text off the machine, and OpenAI's API terms and privacy policy answered 403 to us today so were not read (15 of 30). No deprecation policy. RELEASING.md states that breaking changes bump the minor version before 1.0 (8 of 20). No telemetry was found in the source. Requests to api.openai.com carry `safety_identifier` set to `openai-guardrails-python`, which is described in a source docstring and not in the docs, with no switch found (12 of 20). Fix list for a coding agent, everything this grade says the listing lacks, the biggest gain first (14 items): https://www.anchorterminal.com/fixes/openai-guardrails.md (JSON https://www.anchorterminal.com/fixes/openai-guardrails.json) ### What we couldn't check - unchecked: https://openai.com/policies/services-agreement/ and https://openai.com/policies/privacy-policy/ answered HTTP 403 to our reader on 8 October 2026, so the terms and privacy policy that cover the OpenAI API calls the checks make were not read and `provenance.terms` and `provenance.privacy` are left out. - unchecked: https://openai.com/security/disclosure/ answered HTTP 403, so the disclosure policy and any bounty scope for this package were not read. - unchecked: https://guardrails.openai.com/ is drawn by script, so the configuration wizard and whatever terms it links were not read. - Whether `safety_identifier` can be turned off for calls to api.openai.com. No switch was found in the source. - Why no release was published between 15 December 2025 and 21 July 2026. - The tag v0.3.1 exists in the repository and no 0.3.1 is on PyPI. ### Sources - repository, README, licence and source tree (shallow clone): (seen 2026-10-08) - changelog: (seen 2026-10-08) - release notes: (seen 2026-10-08) - security policy: (seen 2026-10-08) - CI workflow and runs on main: (seen 2026-10-08) - issues and pull requests: (seen 2026-10-08) - contribution policy: (seen 2026-10-08) - release process: (seen 2026-10-08) - quickstart, with error handling modes: (seen 2026-10-08) - streaming and blocking behaviour: (seen 2026-10-08) - jailbreak check and benchmark: (seen 2026-10-08) - prompt injection detection check: (seen 2026-10-08) - Contains PII check: (seen 2026-10-08) - safety identifier source: (seen 2026-10-08) - PyPI package and release dates: (seen 2026-10-08) - PyPI downloads: (seen 2026-10-08) - npm package: (seen 2026-10-08) - TypeScript repository: (seen 2026-10-08) - security.txt: (seen 2026-10-08) ## Who's behind it (provenance 87/100, checked 2026-10-08) | Check | Finding | Points | | --- | --- | --- | | Legal entity named | OpenAI (as named in the LICENSE copyright line and the PyPI author field) | 20/20 | | Domain age | openai.com, registered 2007-01-19 (19 years) | 15/15 | | Endpoint on the vendor's domain | no hosted endpoint | n/a | | Terms of service | nothing hosted, so the MIT licence stands in | 10/10 | | Privacy policy | nothing hosted, not scored | n/a | | Status page | not found | 0/10 | | Changelog | published | 10/10 | | security.txt | valid | 10/10 | A library, not a service. The code is on github.com under the openai organisation, the docs on openai.github.io and the configuration wizard on guardrails.openai.com. The library is governed by the MIT licence in the repository. The model and moderation calls it makes are OpenAI API calls on the user's own key, under OpenAI's API terms. https://openai.com/policies/services-agreement/ and https://openai.com/policies/privacy-policy/ answered HTTP 403 to us on 8 October 2026, so the terms and privacy fields are left out as unread. https://openai.com/.well-known/security.txt is PGP-signed and lists a Bugcrowd contact, a disclosure address and a policy link. It has no Expires line. The copy at cdn.openai.com/security.txt that SECURITY.md links carries an Expires date of 17 January 2024. No status page applies to the library. The OpenAI API it calls has one at https://status.openai.com. RDAP gives 19 January 2007 as the registration date of openai.com. The repository was created on 9 April 2025 and the first PyPI release is dated 6 October 2025. ### Terms and privacy, as read A reading by a fixed set of rules, each answered with the vendor's own sentence. Not legal advice. **Terms of service**. Nothing is hosted by the vendor, so there are no terms of service to read. The MIT licence stands in and the check scores in full. **Privacy policy**. Nothing is hosted by the vendor, so there is no privacy policy to read and the check isn't scored. ## Probe metrics Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. Live uptime, where we poll the endpoint, is under Live and doesn't change the score. ## Strengths - MIT licence, source on GitHub, and three PyPI releases in the 90 days to 8 October 2026 (0.3.0, 0.3.2, 0.3.3) - Twelve built-in checks set in one versioned JSON file across pre-flight, input and output stages - `GuardrailAgent` runs the prompt injection check before and after every tool call in the OpenAI Agents SDK - CI runs ruff, mypy, pyright and tests on Python 3.11 to 3.14, with CodeQL, Dependabot and SHA-pinned actions - The jailbreak page publishes ROC AUC, precision, recall and latency per model on a 4,000-conversation sample ## Weaknesses - By default a check that fails to run returns `tripwire_triggered=False`, so the request continues. Strict mode is opt-in - The README titles the package a preview, the version is 0.3.3, and no release was published between 15 December 2025 and 21 July 2026 - With `stream=True` the output checks run alongside the stream, and the docs say violating content may appear briefly - The docs' own table gives the default jailbreak model, `gpt-4.1-mini`, a recall of 0.000 at a 1 per cent false positive rate - LLM-based checks add a billed model call each. The docs list a median of 1,538 ms for `gpt-4.1-mini` on the jailbreak check - Pull requests from non-collaborators are not accepted, and CHANGELOG.md starts at 0.3.3 ## Before you call it (notes for agents) 1. Pass `raise_guardrail_errors=True` to the client. The default treats a check that failed to run as passed 2. Run `python -m spacy download en_core_web_sm` before using Contains PII, or client initialisation fails 3. Catch `GuardrailTripwireTriggered`, and append a user message to history only after the call returns without it 4. Use `block=true` for Contains PII in the output stage. Masking works only in the pre-flight stage 5. Keep `stream=False` where output must be checked before it is shown, and budget one extra model call per LLM-based check ## Get started Install: ```bash pip install openai-guardrails python -m spacy download en_core_web_sm # only for Contains PII ``` ## Similar tools Ranked by shared capabilities, then score. Same-category tools with no shared capability key are listed last. | Tool | Grade | Score | Rank | Shared capabilities | x402 | Markdown | | --- | --- | --- | --- | --- | --- | --- | | NVIDIA NeMo Guardrails | B | 68.4 | 202 | guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host | no | https://www.anchorterminal.com/tools/nemo-guardrails.md | | Lakera Guard (Check Point AI Guardrails) | C | 59.6 | 492 | guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host | no | https://www.anchorterminal.com/tools/lakera-guard.md | | Guardrails AI | D | 49.6 | 701 | guard.injection, guard.pii, guard.moderation, guard.policy, guard.self-host | no | https://www.anchorterminal.com/tools/guardrails-ai.md | | Google Cloud Model Armor | BB | 77.9 | 16 | guard.injection, guard.pii, guard.moderation, guard.policy | no | https://www.anchorterminal.com/tools/google-model-armor.md | | Amazon Bedrock Guardrails | BB | 74.8 | 62 | guard.injection, guard.pii, guard.moderation, guard.policy | no | https://www.anchorterminal.com/tools/amazon-bedrock-guardrails.md | | Vela 2.0 | B | 66.5 | 268 | guard.pii, guard.injection, guard.moderation, guard.self-host | no | https://www.anchorterminal.com/tools/vela.md | ## Panel reviews (0) Reviewed by the Anchor panel (https://www.anchorterminal.com/reviewers/index.md): . Desk reviews, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. For a desk review, the outcome says whether the reviewer's questions could be answered from public material: success, partial or failure. How reviews work: https://www.anchorterminal.com/reviews/how-it-works.md ## Notable - The README is titled OpenAI Guardrails: Python (Preview). The latest release is 0.3.3 on 10 September 2026 (source: ) - Twelve built-in checks. Keyword Filter, Competitors, Moderation, URL Filter, Secret Keys, Contains PII, Hallucination Detection, Jailbreak, Prompt Injection Detection, NSFW Text, Off Topic Prompts and Custom Prompt Check (source: ) - Default error handling is documented as fail-safe mode. A check that fails to run, for example on an invalid model name, returns `tripwire_triggered=False`. `raise_guardrail_errors=True` raises instead (source: ) - With `stream=True`, output checks run in parallel with the stream and the docs state that violating content may briefly appear before a check triggers (source: ) - The jailbreak page reports a 4,000-conversation benchmark. `gpt-4.1` scores ROC AUC 0.999 and the default `gpt-4.1-mini` 0.928, with recall of 0.000 at a 1 per cent false positive rate and a median time to complete of 1,538 ms (source: ) - Contains PII uses Presidio with a spaCy model the user installs, adds CVV and BIC/SWIFT recognisers, and can look inside Base64, URL-encoded and hex strings (source: ) - Calls to api.openai.com carry `safety_identifier` set to `openai-guardrails-python`. Azure and custom endpoints do not receive it (source: ) - Pull requests are limited to repository collaborators. Others are asked to open issues (source: ) - A TypeScript package, `@openai/guardrails` 0.3.0, is published from a separate repository and needs Node.js 22.13 or later (source: ) - `openai-guardrails` had 104,530 PyPI downloads in the week to 8 October 2026 (source: ) ## Compare - [Amazon Bedrock Guardrails vs OpenAI Guardrails](https://www.anchorterminal.com/compare/amazon-bedrock-guardrails-vs-openai-guardrails.md): BB 74.8 vs B 69.5 - [Azure AI Content Safety (Prompt Shields) vs OpenAI Guardrails](https://www.anchorterminal.com/compare/azure-ai-content-safety-vs-openai-guardrails.md): C 60.7 vs B 69.5 - [Cisco AI Defense Inspection API vs OpenAI Guardrails](https://www.anchorterminal.com/compare/cisco-ai-defense-inspection-vs-openai-guardrails.md): C 61.3 vs B 69.5 - [Google Cloud Model Armor vs OpenAI Guardrails](https://www.anchorterminal.com/compare/google-model-armor-vs-openai-guardrails.md): BB 77.9 vs B 69.5 - [Guardrails AI vs OpenAI Guardrails](https://www.anchorterminal.com/compare/guardrails-ai-vs-openai-guardrails.md): D 49.6 vs B 69.5 - [Lakera Guard (Check Point AI Guardrails) vs OpenAI Guardrails](https://www.anchorterminal.com/compare/lakera-guard-vs-openai-guardrails.md): C 59.6 vs B 69.5 - [LlamaFirewall vs OpenAI Guardrails](https://www.anchorterminal.com/compare/llamafirewall-vs-openai-guardrails.md): D 50.8 vs B 69.5 - [NVIDIA NeMo Guardrails vs OpenAI Guardrails](https://www.anchorterminal.com/compare/nemo-guardrails-vs-openai-guardrails.md): B 68.4 vs B 69.5 - [OpenAI Guardrails vs Prisma AIRS AI Runtime Security API](https://www.anchorterminal.com/compare/openai-guardrails-vs-prisma-airs.md): B 69.5 vs B 62.8 - [Granite Guardian vs OpenAI Guardrails](https://www.anchorterminal.com/compare/granite-guardian-vs-openai-guardrails.md): C 60.1 vs B 69.5 - [Llama Guard 4 vs OpenAI Guardrails](https://www.anchorterminal.com/compare/llama-guard-vs-openai-guardrails.md): D 49.1 vs B 69.5 - [Mistral Moderation API vs OpenAI Guardrails](https://www.anchorterminal.com/compare/mistral-moderation-vs-openai-guardrails.md): C 58.4 vs B 69.5 - [OpenAI Guardrails vs OpenAI Moderation API](https://www.anchorterminal.com/compare/openai-guardrails-vs-openai-moderation.md): B 69.5 vs BB 71.3 - [Presidio vs OpenAI Guardrails](https://www.anchorterminal.com/compare/microsoft-presidio-vs-openai-guardrails.md): B 66 vs B 69.5 ## Verify this listing For the vendor. The badge or a plain link to this page verifies the listing, from a page on openai.com or one of its subdomains, or the README of github.com/openai/openai-guardrails-python. It shows the listing is the vendor's and that the vendor knows it's here, and it never changes a grade, rank or review. The vendor sends the page's address to `POST https://www.anchorterminal.com/api/v1/verify` as `{"slug": "openai-guardrails", "url": "…"}`, or calls the `verify_listing` tool at https://www.anchorterminal.com/mcp. We fetch the page once, then again every week; two failed checks in a row and the verification lapses, and a later pass restores it. What we check: https://www.anchorterminal.com/builders/index.md#verify HTML badge: ```html OpenAI Guardrails on Anchor Terminal ``` Markdown badge, for a README: ```markdown [![OpenAI Guardrails on Anchor Terminal](https://www.anchorterminal.com/badges/openai-guardrails.svg)](https://www.anchorterminal.com/tools/openai-guardrails) ``` Plain link: ```html OpenAI Guardrails on Anchor Terminal ``` ## Share this listing For the vendor. Sharing assets for social media, two PNGs of 1200 × 630 that say OpenAI Guardrails is listed on Anchor Terminal, with the vendor's logo and this page's address and no grade or score. - Dark: https://www.anchorterminal.com/assets/share/openai-guardrails-dark.png - Light: https://www.anchorterminal.com/assets/share/openai-guardrails-light.png