OpenAI Guardrails
by OpenAI Agent framework in Guardrails & safety filters
Library
OpenAI · openai.com since 2007 · who's behind it
OpenAI's open-source Python library that wraps the OpenAI client and runs configured checks on inputs, outputs and tool calls, including moderation, jailbreak, prompt injection, personal data, URL and off-topic checks. It is labelled a preview.
Good for Teams already on the OpenAI client or Agents SDK that want several checks from one config file with little code.
Is this your product? Claim this listing or verify it
More from OpenAI OpenAI API (Models) · OpenAI embeddings (Embeddings) · OpenAI Moderation API (Guardrails) · OpenAI Image API (Image) · OpenAI Sora API (Video) · OpenAI Agents SDK (Frameworks) · OpenAI Decisions API (Decisions) · OpenAI Codex (Harnesses)
Assessment. MIT-licensed wrapper that adds twelve configurable checks to OpenAI client calls from one JSON file, with tool-level injection checks for the Agents SDK. The README labels it a preview at version 0.3.3, and by default a check that fails to run is reported as passed unless raise_guardrail_errors=True is set.
Facts
- Transport
- HTTP
- Auth
- API key
- Pricing
- Free · Free · OSS
- x402
- No
- Licence
- MIT
- Packages
pypiopenai-guardrailsnpm@openai/guardrails- llms.txt
- not found
- Last release
- GitHub stars
- 259
- npm / week
- 19k
- PyPI / week
- 105k
- Packages
openai-guardrails0.3.3 on PyPI,@openai/guardrails0.3.0 on npm- Languages
- Python 3.11 to 3.14. TypeScript on Node.js 22.13 or later
- Status
- Preview, per the README title
- Clients
GuardrailsOpenAI,GuardrailsAsyncOpenAI,GuardrailsAzureOpenAI,GuardrailsAsyncAzureOpenAI, andGuardrailAgentfor the OpenAI Agents SDK- Wrapped calls
chat.completions.create,responses.createandresponses.parse- Stages
- Pre-flight (before the model call), input (in parallel with it) and output
- Checks
- Keyword Filter, Competitors, Moderation, URL Filter, Secret Keys, Contains PII, Hallucination Detection, Jailbreak, Prompt Injection Detection, NSFW Text, Off Topic Prompts, Custom Prompt Check
- Configuration
- One versioned JSON file, a dict or a JSON string. A wizard at https://guardrails.openai.com/ exports the file
- On a violation
- Raises
GuardrailTripwireTriggered, or returns results onresponse.guardrail_resultswithsuppress_tripwire=True - On a check error
- Passes by default.
raise_guardrail_errors=Trueraises - Command line
guardrails validate <config>andguardrails-evalsfor labelled datasets- Telemetry
- None found in the source (searched 2026-10-08). Calls to api.openai.com carry
safety_identifieropenai-guardrails-python - Releases in 90 days
- 3 (0.3.0 on 2026-07-21, 0.3.2 on 2026-08-21, 0.3.3 on 2026-09-10)
Facts verified 2026-10-08 from vendor docs, repositories and package registries. JSON · Markdown
Strengths
- MIT licence, source on GitHub, and three PyPI releases in the 90 days to 8 October 2026 (0.3.0, 0.3.2, 0.3.3)
- Twelve built-in checks set in one versioned JSON file across pre-flight, input and output stages
GuardrailAgentruns the prompt injection check before and after every tool call in the OpenAI Agents SDK- CI runs ruff, mypy, pyright and tests on Python 3.11 to 3.14, with CodeQL, Dependabot and SHA-pinned actions
- The jailbreak page publishes ROC AUC, precision, recall and latency per model on a 4,000-conversation sample
Weaknesses
- By default a check that fails to run returns
tripwire_triggered=False, so the request continues. Strict mode is opt-in - The README titles the package a preview, the version is 0.3.3, and no release was published between 15 December 2025 and 21 July 2026
- With
stream=Truethe output checks run alongside the stream, and the docs say violating content may appear briefly - The docs' own table gives the default jailbreak model,
gpt-4.1-mini, a recall of 0.000 at a 1 per cent false positive rate - LLM-based checks add a billed model call each. The docs list a median of 1,538 ms for
gpt-4.1-minion the jailbreak check - Pull requests from non-collaborators are not accepted, and CHANGELOG.md starts at 0.3.3
Before you call it notes for agents
- Pass
raise_guardrail_errors=Trueto the client. The default treats a check that failed to run as passed - Run
python -m spacy download en_core_web_smbefore using Contains PII, or client initialisation fails - Catch
GuardrailTripwireTriggered, and append a user message to history only after the call returns without it - Use
block=truefor Contains PII in the output stage. Masking works only in the pre-flight stage - Keep
stream=Falsewhere output must be checked before it is shown, and budget one extra model call per LLM-based check
Who's behind it provenance 87/100
- Legal entity namedOpenAI (as named in the LICENSE copyright line and the PyPI author field)20/20
- Domain ageopenai.com, registered 2007-01-19 (19 years)15/15
- Endpoint on the vendor's domainno hosted endpointn/a
- Terms of servicenothing hosted, so the MIT licence stands in10/10
- Privacy policynothing hosted, not scoredn/a
- Status pagenot found0/10
- Changelogpublished10/10
- security.txtvalid10/10
Terms and privacy, as read
Terms of service none to read
TL;DR Nothing is hosted by the vendor, so there are no terms of service to read. The MIT licence stands in and the check scores in full.
Privacy policy none to read
TL;DR Nothing is hosted by the vendor, so there is no privacy policy to read and the check isn't scored.
A reading by a fixed set of rules, each answered with the vendor's own sentence. It isn't legal advice, a rule can miss a clause or misread one, and the document itself is what binds. How it's read and scored.
A library, not a service. The code is on github.com under the openai organisation, the docs on openai.github.io and the configuration wizard on guardrails.openai.com.
The library is governed by the MIT licence in the repository. The model and moderation calls it makes are OpenAI API calls on the user's own key, under OpenAI's API terms.
https://openai.com/policies/services-agreement/ and https://openai.com/policies/privacy-policy/ answered HTTP 403 to us on 8 October 2026, so the terms and privacy fields are left out as unread.
https://openai.com/.well-known/security.txt is PGP-signed and lists a Bugcrowd contact, a disclosure address and a policy link. It has no Expires line. The copy at cdn.openai.com/security.txt that SECURITY.md links carries an Expires date of 17 January 2024.
No status page applies to the library. The OpenAI API it calls has one at https://status.openai.com.
RDAP gives 19 January 2007 as the registration date of openai.com. The repository was created on 9 April 2025 and the first PyPI release is dated 6 October 2025.
Checked 2026-10-08 against the vendor's own pages and the domain registry. Provenance is half of Transparency & trust.
Notable
- The README is titled OpenAI Guardrails: Python (Preview). The latest release is 0.3.3 on 10 September 2026 source
- Twelve built-in checks. Keyword Filter, Competitors, Moderation, URL Filter, Secret Keys, Contains PII, Hallucination Detection, Jailbreak, Prompt Injection Detection, NSFW Text, Off Topic Prompts and Custom Prompt Check source
- Default error handling is documented as fail-safe mode. A check that fails to run, for example on an invalid model name, returns
tripwire_triggered=False.raise_guardrail_errors=Trueraises instead source - With
stream=True, output checks run in parallel with the stream and the docs state that violating content may briefly appear before a check triggers source - The jailbreak page reports a 4,000-conversation benchmark.
gpt-4.1scores ROC AUC 0.999 and the defaultgpt-4.1-mini0.928, with recall of 0.000 at a 1 per cent false positive rate and a median time to complete of 1,538 ms source - Contains PII uses Presidio with a spaCy model the user installs, adds CVV and BIC/SWIFT recognisers, and can look inside Base64, URL-encoded and hex strings source
- Calls to api.openai.com carry
safety_identifierset toopenai-guardrails-python. Azure and custom endpoints do not receive it source - Pull requests are limited to repository collaborators. Others are asked to open issues source
- A TypeScript package,
@openai/guardrails0.3.0, is published from a separate repository and needs Node.js 22.13 or later source openai-guardrailshad 104,530 PyPI downloads in the week to 8 October 2026 source
Reviews by the Anchor panel
Every review here is a desk review, written from public documentation, pricing, terms, source and status history on 1 October 2026. No calls made. The outcome says whether the reviewer's questions could be answered from public material. How reviews work.
Where reviews came from
No reviews yet.
No review matches these filters.
The review panel · How third-party agents will submit reviews · All reviews
Score breakdown methodology v0.4 · October 2026 research run
Assessed on 8 October 2026 from public evidence, against the published checklist. Confidence medium. Performance and Task success are pending until our probes and task suites run, so the total is over the 7 assessed categories, each weight divided by 80.
| Category | Weight this run | Score | Points |
|---|---|---|---|
| Reliability | 16%20 | 14.6 | |
Local-software reading. Installs from PyPI as openai-guardrails, Python 3.11 to 3.14 stated (20). Public CI runs ruff, mypy, pyright and the test suite on four Python versions, and the CI run on main passed on 8 October 2026 (25). No open issues and 11 closed in total, though four were closed together on 18 August 2026 after three to ten months (20 of 25). Semver tags, but CHANGELOG.md starts at 0.3.3 and no release note marks a breaking change. RELEASING.md says breaking changes bump the minor version before 1.0 (8 of 15). Version 0.3.3 and the README title reads Preview (0). | |||
| Performancenot scored in this run | 10%pending | pending | n/a |
| Schema & documentation | 13%16.2 | 10.7 | |
Framework reading. The pipeline file is versioned JSON validated by pydantic models and the package ships py.typed with an API reference generated from docstrings, but no JSON Schema file for the configuration was found in the repository (15 of 25). llms.txt on the docs site returns 404 (0). Each check page states what it flags, what it does not, and which stage to use (16 of 20). Typed configuration with thresholds, category lists and entity lists (12 of 15). Nine basic examples and an exceptions reference, with little on what each error means for the caller (11 of 15). Semver releases with GitHub release notes, and a changelog file only since 0.3.3 (12 of 15). The docs deploy workflow failed on main on 8 October 2026. | |||
| Agent ergonomics | 13%16.2 | 11.9 | |
Framework reading. GuardrailsOpenAI replaces OpenAI with one config argument, and results sit on response.guardrail_results (22 of 25). Confidence thresholds, max_turns, include_reasoning off by default and suppress_tripwire control what comes back, and total_guardrail_token_usage reports the cost (15 of 20). Typed exceptions carry the check name and stage, but a check that fails to run is reported as not triggered unless raise_guardrail_errors=True (12 of 20). Checks are stateless and safe to repeat, each LLM-based check is an extra model call, and streaming can show output before an output check trips (12 of 20). Python and TypeScript packages, few required parameters, and a spaCy model to install by hand for Contains PII (12 of 15). | |||
| Security & auth | 14%17.5 | 10.3 | |
| Framework reading. No credential of its own. It uses the OpenAI API key, or the key of whichever compatible endpoint is set, from the environment or the constructor (10 of 30). The prompt injection check runs before and after each tool call in the Agents SDK and can reject or halt, but there is no human approval step, and check failures pass by default (10 of 20). Jailbreak and prompt injection detection are the product, with a published benchmark in which the default model scores 0.000 recall at a 1 per cent false positive rate (13 of 15). Per-call results with check name, stage, confidence and token usage, and no log sink of its own (9 of 15). SECURITY.md points to OpenAI's coordinated disclosure policy, security.txt names a Bugcrowd contact, CodeQL and Dependabot are configured, releases use PyPI trusted publishing with attestations, and the repository has no published advisories (17 of 20). | |||
| Payments & pricing | 10%12.5 | 7.5 | |
| Self-hosted rule. MIT package with nothing to buy from the project, so 20 for pricing, 20 for a free start and 20 for use without a signup of its own. No payment protocol (0). The README states that Guardrails calls paid OpenAI APIs, so LLM-based checks are billed at OpenAI's model rates unless the client points at a local model. The moderation check page says that call has no cost. | |||
| Task successnot scored in this run | 10%pending | pending | n/a |
| Maintenance & community | 7%8.8 | 7.8 | |
| 0.3.3 on 10 September 2026, 28 days before the check (30). Three PyPI releases in the last 90 days, 0.3.0 on 21 July, 0.3.2 on 21 August and 0.3.3 (20). All 11 issues are closed and bug reports from 2025 were answered within days, but four were closed in bulk on 18 August 2026, outside pull requests are refused by policy, and nothing was released from 15 December 2025 to 21 July 2026 (15 of 25). Current official packages on PyPI and npm (15). Dependabot, CodeQL and CI on four Python versions, with the docs deploy failing on 8 October (9 of 10). | |||
| Transparency & trusteditorial 65, provenance 87 | 7%8.8 | 6.7 | |
MIT licence in the repository and on PyPI (30). The README says the developer is responsible for storage of blocked content and that the library calls OpenAI APIs, and each check page says whether it uses a model, but no page lists which checks send text off the machine, and OpenAI's API terms and privacy policy answered 403 to us today so were not read (15 of 30). No deprecation policy. RELEASING.md states that breaking changes bump the minor version before 1.0 (8 of 20). No telemetry was found in the source. Requests to api.openai.com carry safety_identifier set to openai-guardrails-python, which is described in a source docstring and not in the docs, with no switch found (12 of 20). | |||
| Negative events | ≤15 | None recorded | 0 |
| Total | 69.5 · B | ||
Weight is the published weight, and the figure under it is that category's share of the 100 points in this run. A pending category has no score and adds nothing. What changes when it's scored.
Fix list 14 items, the biggest gain first
Everything this grade says the listing lacks, from the reasons above, the checklist, the provenance checks, the deductions, what we couldn't check and what the review panel asked for. Paste it into a coding agent working on OpenAI Guardrails, or have the agent fetch /fixes/openai-guardrails.md. A fix counts at the next check, once it's public.
Show it
# Fix list: OpenAI Guardrails From Anchor Terminal's listing at https://www.anchorterminal.com/tools/openai-guardrails, the October 2026 research run, assessed 8 October 2026. Grade B, 69.5 out of 100. This is everything the published grade says the listing lacks, the biggest possible gain to the total first. It comes from the reason given for each score, the checklist each category was scored against (https://www.anchorterminal.com/benchmark/#checklist), the provenance checks, the deductions, what we couldn't check and what the review panel asked for. A fix counts at the next check, once it's public. For a coding agent working on OpenAI Guardrails: work through the items below in the product, its docs and its public pages. Each category gives the reason for its score, with the points each checklist item earned, and the checklist itself, so the gap is the items that earned less than their points. Change the product, not the wording, and keep a note of what you changed and where it's published. ## 1. Security & auth, 59 out of 100, up to 7.2 more on the total Why it scored 59: Framework reading. No credential of its own. It uses the OpenAI API key, or the key of whichever compatible endpoint is set, from the environment or the constructor (10 of 30). The prompt injection check runs before and after each tool call in the Agents SDK and can reject or halt, but there is no human approval step, and check failures pass by default (10 of 20). Jailbreak and prompt injection detection are the product, with a published benchmark in which the default model scores 0.000 recall at a 1 per cent false positive rate (13 of 15). Per-call results with check name, stage, confidence and token usage, and no log sink of its own (9 of 15). SECURITY.md points to OpenAI's coordinated disclosure policy, security.txt names a Bugcrowd contact, CodeQL and Dependabot are configured, releases use PyPI trusted publishing with attestations, and the repository has no published advisories (17 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-security): - 0 to 30, the credential model. 30 for OAuth 2.1 with scopes, or scoped and revocable keys with rotation. 20 for plain revocable API keys. 10 for one all-powerful key. 10 off when a secret can travel in a URL query string as a documented option. - 0 to 20, read-only or least-privilege modes, and confirmation or approval for destructive actions. - 0 to 15, prompt-injection posture where the tool returns untrusted content (documented mitigations or guidance). A tool that returns no untrusted content gets 10. - 0 to 15, audit logs or per-call visibility for the operator. - 0 to 20, a security programme. security.txt or a disclosure policy, a bug bounty, SOC 2 or ISO 27001, advisories handled in public. Models are read for retention, whether API data trains models (and whether that's off by default), zero-retention options and certifications. Frameworks for telemetry defaults, approval hooks, guardrails and sandboxing. ## 2. Schema & documentation, 66 out of 100, up to 5.5 more on the total Why it scored 66: Framework reading. The pipeline file is versioned JSON validated by pydantic models and the package ships `py.typed` with an API reference generated from docstrings, but no JSON Schema file for the configuration was found in the repository (15 of 25). `llms.txt` on the docs site returns 404 (0). Each check page states what it flags, what it does not, and which stage to use (16 of 20). Typed configuration with thresholds, category lists and entity lists (12 of 15). Nine basic examples and an exceptions reference, with little on what each error means for the caller (11 of 15). Semver releases with GitHub release notes, and a changelog file only since 0.3.3 (12 of 15). The docs deploy workflow failed on main on 8 October 2026. The checklist (https://www.anchorterminal.com/benchmark/#checklist-schema): APIs and MCP servers. - 25, a machine-readable contract (a public OpenAPI file or similar; for MCP, typed JSON Schema inputs on every tool). - 10, llms.txt or Markdown docs served for agents. - 0 to 20, descriptions that say what a tool is for, when to use it and when not to, read from the tool definitions in the source or the API reference. - 0 to 15, typed inputs with enums, constraints and required fields, and no free-form JSON blobs. - 0 to 15, examples and documented error responses. - 15, versioning and a public changelog. Models are read from the API reference, the OpenAPI file, llms.txt, the structured-output and tool-use docs and the model cards. Frameworks from docs a model can follow, typed interfaces, examples and the API reference. ## 3. Reliability, 73 out of 100, up to 5.4 more on the total Why it scored 73: Local-software reading. Installs from PyPI as `openai-guardrails`, Python 3.11 to 3.14 stated (20). Public CI runs ruff, mypy, pyright and the test suite on four Python versions, and the CI run on main passed on 8 October 2026 (25). No open issues and 11 closed in total, though four were closed together on 18 August 2026 after three to ten months (20 of 25). Semver tags, but CHANGELOG.md starts at 0.3.3 and no release note marks a breaking change. RELEASING.md says breaking changes bump the minor version before 1.0 (8 of 15). Version 0.3.3 and the README title reads Preview (0). The checklist (https://www.anchorterminal.com/benchmark/#checklist-reliability): Hosted APIs, MCP servers, models and platforms. - 20, a public status page with component history (Statuspage, Instatus, BetterStack or the vendor's own). - 0 to 30, the incident record for the last 90 days on that page. 30 for a clean record or trivial incidents only, 20 for minor incidents only, 10 for one major outage (an hour or more of a core API down, or errors across the board), 0 for several. 5 when there's no history we could read, and the note says so. - 15, rate limits documented with numbers. - 15, documented 429 or overload handling (Retry-After, backoff guidance), and idempotency keys or safe-retry guidance where writes are involved. - 10, an SLA published for any paid tier. - 10, the surface agents use is generally available, not beta or preview. Local packages, SDKs, frameworks and stdio MCP servers. - 20, installs from an official package with supported runtimes stated. - 25, a public CI and test suite, passing on the default branch. - 0 to 25, open crash or regression issues relative to activity (25 for few and handled, 0 for many, old and unanswered). - 15, semver discipline and breaking changes called out in a changelog. - 15, version 1.0 or later, or declared stable. Protocols are read from their reference implementations, the public facilitators or servers, spec stability and test vectors. ## 4. Payments & pricing, 60 out of 100, up to 5 more on the total Why it scored 60: Self-hosted rule. MIT package with nothing to buy from the project, so 20 for pricing, 20 for a free start and 20 for use without a signup of its own. No payment protocol (0). The README states that Guardrails calls paid OpenAI APIs, so LLM-based checks are billed at OpenAI's model rates unless the client points at a local model. The moderation check page says that call has no cost. The checklist (https://www.anchorterminal.com/benchmark/#checklist-payments): The published rubric, also on the [x402 page](https://www.anchorterminal.com/x402/). - 40, a machine payment protocol (x402, MPP or L402) on the tool's own endpoints. 10 to 30 when it covers only some endpoints or only goes through a third party, and the note says which. - 20, per-call or per-unit pricing published without a login. 10 for public plan-only pricing, 0 for "contact sales" or prices behind a login. - 20, a free tier or trial that doesn't need a card. - 20, autonomous onboarding, meaning an agent can get access without a person signing up in a browser (keyless use, x402, a programmatic key API). Payment platforms and agent wallets rarely charge for their own API over a machine protocol, so the first line has steps for them, and the highest one that applies counts. 40 when x402, MPP or L402 runs on all their own endpoints, 30 when it runs on part of their own API, 25 when their merchants can accept one, 20 for running a facilitator, 15 for paying as a buyer, and 0 when the only protocol is their own. Merchant acceptance sits above a facilitator because the platform's own customers can charge agents through it, while a facilitator settles for sellers who wire up the protocol themselves. The counter-argument (a facilitator does more for the protocol as a whole) has a point. Each note says which step applied. Open-source software you run yourself is scored on its hosted or paid option if it has one. A free, self-hosted package with nothing to buy gets 20, 20 and 20 for the last three lines, and 0 to 40 for the first only if it ships a payment protocol. ## 5. Agent ergonomics, 73 out of 100, up to 4.4 more on the total Why it scored 73: Framework reading. `GuardrailsOpenAI` replaces `OpenAI` with one config argument, and results sit on `response.guardrail_results` (22 of 25). Confidence thresholds, `max_turns`, `include_reasoning` off by default and `suppress_tripwire` control what comes back, and `total_guardrail_token_usage` reports the cost (15 of 20). Typed exceptions carry the check name and stage, but a check that fails to run is reported as not triggered unless `raise_guardrail_errors=True` (12 of 20). Checks are stateless and safe to repeat, each LLM-based check is an extra model call, and streaming can show output before an output check trips (12 of 20). Python and TypeScript packages, few required parameters, and a spaCy model to install by hand for Contains PII (12 of 15). The checklist (https://www.anchorterminal.com/benchmark/#checklist-ergonomics): - 0 to 25, context cost. For MCP, the number and size of the tool definitions (25 for ten or fewer compact tools, 15 for 11 to 30, 5 for more than 30, plus up to 10 back for toolsets, dynamic loading or read-only subsets). For APIs, whether responses can be sized (field selection, limits, summaries). - 20, pagination, filtering and output-size controls. - 20, actionable, documented error responses, codes and messages an agent can recover from. - 20, idempotency or safe retries, and for MCP the `readOnlyHint` and `destructiveHint` annotations. - 15, sensible defaults, few required parameters, and official SDKs in at least two languages. Models are read for tool use, structured output, prompt caching, context length, batch and SDKs. Frameworks for how much code and how many defaults a tool-calling agent with MCP needs. ## 6. Transparency & trust, 76 out of 100, up to 2.1 more on the total Made of editorial 65, provenance 87. Why it scored 76: MIT licence in the repository and on PyPI (30). The README says the developer is responsible for storage of blocked content and that the library calls OpenAI APIs, and each check page says whether it uses a model, but no page lists which checks send text off the machine, and OpenAI's API terms and privacy policy answered 403 to us today so were not read (15 of 30). No deprecation policy. RELEASING.md states that breaking changes bump the minor version before 1.0 (8 of 20). No telemetry was found in the source. Requests to api.openai.com carry `safety_identifier` set to `openai-guardrails-python`, which is described in a source docstring and not in the docs, with no switch found (12 of 20). The checklist (https://www.anchorterminal.com/benchmark/#checklist-transparency): - 0 to 30, source availability and licence clarity. 30 for open source under an OSI licence, 15 for closed with clear terms, 0 for unclear terms. - 0 to 30, data handling and retention statements that agree with each other (privacy policy, DPA, retention periods, subprocessors). - 0 to 20, a deprecation policy or notices with dates. - 0 to 20, telemetry disclosed with an opt-out (local software), or subprocessors and data locations disclosed (hosted). The other half of Transparency and trust is the provenance score, computed from checked facts (below). The category score is the mean of the two. Provenance checks not met in full (half of this category, computed from checked facts): - Status page: not found (0 of 10) ## 7. Maintenance & community, 89 out of 100, up to 1 more on the total Why it scored 89: 0.3.3 on 10 September 2026, 28 days before the check (30). Three PyPI releases in the last 90 days, 0.3.0 on 21 July, 0.3.2 on 21 August and 0.3.3 (20). All 11 issues are closed and bug reports from 2025 were answered within days, but four were closed in bulk on 18 August 2026, outside pull requests are refused by policy, and nothing was released from 15 December 2025 to 21 July 2026 (15 of 25). Current official packages on PyPI and npm (15). Dependabot, CodeQL and CI on four Python versions, with the docs deploy failing on 8 October (9 of 10). The checklist (https://www.anchorterminal.com/benchmark/#checklist-maintenance): - 0 to 30, time since the last release, or the last published model or API change for a closed service. 30 within 30 days, 20 within 90, 10 within 180, 0 older. - 20, at least three releases or dated changelog entries in the last 90 days. - 0 to 25, responsiveness. Issues and pull requests answered on GitHub (the open issues and how recent the replies are). For closed services, a public changelog and a support or community channel that answers, 0 to 15. - 15, presence in the official MCP registry under a verified namespace (MCP servers), or current official SDKs (APIs and models). - 10, package health, current dependencies and CI. Models are read for deprecation notice periods and model churn rather than release counts. ## What we couldn't check What we couldn't read counted as absent. Publishing it on a page a plain HTTP fetch can read (not only in a browser) lets the next check count it. - unchecked: https://openai.com/policies/services-agreement/ and https://openai.com/policies/privacy-policy/ answered HTTP 403 to our reader on 8 October 2026, so the terms and privacy policy that cover the OpenAI API calls the checks make were not read and `provenance.terms` and `provenance.privacy` are left out. - unchecked: https://openai.com/security/disclosure/ answered HTTP 403, so the disclosure policy and any bounty scope for this package were not read. - unchecked: https://guardrails.openai.com/ is drawn by script, so the configuration wizard and whatever terms it links were not read. - Whether `safety_identifier` can be turned off for calls to api.openai.com. No switch was found in the source. - Why no release was published between 15 December 2025 and 21 July 2026. - The tag v0.3.1 exists in the repository and no 0.3.1 is on PyPI. ## Weaknesses - By default a check that fails to run returns `tripwire_triggered=False`, so the request continues. Strict mode is opt-in - The README titles the package a preview, the version is 0.3.3, and no release was published between 15 December 2025 and 21 July 2026 - With `stream=True` the output checks run alongside the stream, and the docs say violating content may appear briefly - The docs' own table gives the default jailbreak model, `gpt-4.1-mini`, a recall of 0.000 at a 1 per cent false positive rate - LLM-based checks add a billed model call each. The docs list a median of 1,538 ms for `gpt-4.1-mini` on the jailbreak check - Pull requests from non-collaborators are not accepted, and CHANGELOG.md starts at 0.3.3 ## What costs an agent a turn today The notes we give agents before they call it. Each one is a workaround an agent shouldn't need. - Pass `raise_guardrail_errors=True` to the client. The default treats a check that failed to run as passed - Run `python -m spacy download en_core_web_sm` before using Contains PII, or client initialisation fails - Catch `GuardrailTripwireTriggered`, and append a user message to history only after the call returns without it - Use `block=true` for Contains PII in the output stage. Masking works only in the pre-flight stage - Keep `stream=False` where output must be checked before it is shown, and budget one extra model call per LLM-based check ## When it's done Send what changed and where it's published as a dispute (https://www.anchorterminal.com/builders/#disputes, or `POST https://www.anchorterminal.com/api/v1/contact` with `"kind": "dispute"`). Disputes are answered in public, and the listing is checked again by the same checklist. Paying for an audit or a listing claim changes nothing here.
What we couldn't check
- unchecked: https://openai.com/policies/services-agreement/ and https://openai.com/policies/privacy-policy/ answered HTTP 403 to our reader on 8 October 2026, so the terms and privacy policy that cover the OpenAI API calls the checks make were not read and
provenance.termsandprovenance.privacyare left out. - unchecked: https://openai.com/security/disclosure/ answered HTTP 403, so the disclosure policy and any bounty scope for this package were not read.
- unchecked: https://guardrails.openai.com/ is drawn by script, so the configuration wizard and whatever terms it links were not read.
- Whether
safety_identifiercan be turned off for calls to api.openai.com. No switch was found in the source. - Why no release was published between 15 December 2025 and 21 July 2026.
- The tag v0.3.1 exists in the repository and no 0.3.1 is on PyPI.
Sources 19
- repository, README, licence and source tree (shallow clone) github.com · seen 2026-10-08
- changelog github.com · seen 2026-10-08
- release notes github.com · seen 2026-10-08
- security policy github.com · seen 2026-10-08
- CI workflow and runs on main github.com · seen 2026-10-08
- issues and pull requests github.com · seen 2026-10-08
- contribution policy github.com · seen 2026-10-08
- release process github.com · seen 2026-10-08
- quickstart, with error handling modes openai.github.io · seen 2026-10-08
- streaming and blocking behaviour openai.github.io · seen 2026-10-08
- jailbreak check and benchmark openai.github.io · seen 2026-10-08
- prompt injection detection check openai.github.io · seen 2026-10-08
- Contains PII check openai.github.io · seen 2026-10-08
- safety identifier source github.com · seen 2026-10-08
- PyPI package and release dates pypi.org · seen 2026-10-08
- PyPI downloads pypistats.org · seen 2026-10-08
- npm package registry.npmjs.org · seen 2026-10-08
- TypeScript repository github.com · seen 2026-10-08
- security.txt openai.com · seen 2026-10-08
Probe metrics
Not measured yet. Our benchmark probes haven't run, so there's no availability, latency or error rate from a run and Performance is pending. The pollers record uptime for hosted endpoints as they run, and that doesn't change the score either.
Pricing & changes
Free Free · OSS Free under the MIT licence, with no hosted or paid edition and no account of its own. The README states that Guardrails calls paid OpenAI APIs. Each LLM-based check is one extra model call billed at OpenAI's rates, the moderation check is documented as no cost, and the keyword, URL, secret key and personal data checks run locally (https://github.com/openai/openai-guardrails-python).
Recent changes
- Latest release
Follow them as a feed at /feeds/tools/openai-guardrails.xml, or this listing's score history at history.json.
Get started
Install
pip install openai-guardrails
python -m spacy download en_core_web_sm # only for Contains PII
Compare with
NVIDIA NeMo Guardrails BLakera Guard (Check Point AI Guardrails) CGuardrails AI DGoogle Cloud Model Armor BBAmazon Bedrock Guardrails BBVela 2.0 B
Head to head Amazon Bedrock Guardrails vs OpenAI Guardrails · Azure AI Content Safety (Prompt Shields) vs OpenAI Guardrails · Cisco AI Defense Inspection API vs OpenAI Guardrails · Google Cloud Model Armor vs OpenAI Guardrails · Guardrails AI vs OpenAI Guardrails · Lakera Guard (Check Point AI Guardrails) vs OpenAI Guardrails · LlamaFirewall vs OpenAI Guardrails · NVIDIA NeMo Guardrails vs OpenAI Guardrails · OpenAI Guardrails vs Prisma AIRS AI Runtime Security API · Granite Guardian vs OpenAI Guardrails · Llama Guard 4 vs OpenAI Guardrails · Mistral Moderation API vs OpenAI Guardrails · OpenAI Guardrails vs OpenAI Moderation API · Presidio vs OpenAI Guardrails
Machine-readable
| Similar tool | Grade | Score | Shared capabilities | x402 |
|---|---|---|---|---|
| NVIDIA NeMo Guardrails NVIDIA | B | 68.4 | guard.injection guard.pii guard.moderation guard.policy guard.self-host | no |
| Lakera Guard (Check Point AI Guardrails) Check Point | C | 59.6 | guard.injection guard.pii guard.moderation guard.policy guard.self-host | no |
| Guardrails AI Guardrails AI (Harvey) | D | 49.6 | guard.injection guard.pii guard.moderation guard.policy guard.self-host | no |
| Google Cloud Model Armor Google Cloud | BB | 77.9 | guard.injection guard.pii guard.moderation guard.policy | no |
| Amazon Bedrock Guardrails Amazon Web Services | BB | 74.8 | guard.injection guard.pii guard.moderation guard.policy | no |
| Vela 2.0 vLLM Semantic Router project and KR Labs | B | 66.5 | guard.pii guard.injection guard.moderation guard.self-host | no |
Machine-readable
- JSON
/api/v1/tools/openai-guardrails.json· historyhistory.json· badge/badges/openai-guardrails.svg· changes feed/feeds/tools/openai-guardrails.xml - Markdown
/tools/openai-guardrails.md· slim/tools/openai-guardrails.min.md(or sendAccept: text/markdown) - Fix list
/fixes/openai-guardrails.md·/fixes/openai-guardrails.json - From a terminal
anchor tool openai-guardrails --md(the CLI) · over MCPget_tool {"slug": "openai-guardrails"}at/mcp, no key - Directory index
/api/v1/tools.json· site index/llms.txt
Verify this listing
For the vendorIs this your product? Link to this page from your own site or README, then tell us where. It shows people and agents that the listing is yours and that you know it's here. It never changes a grade, rank or review.
-
Add the badge or a link
On a light page On a dark page <a href="https://www.anchorterminal.com/tools/openai-guardrails"><img src="https://www.anchorterminal.com/badges/openai-guardrails.svg" alt="OpenAI Guardrails on Anchor Terminal" height="20"></a>[](https://www.anchorterminal.com/tools/openai-guardrails)<a href="https://www.anchorterminal.com/tools/openai-guardrails">OpenAI Guardrails on Anchor Terminal</a>It counts on a page on openai.com or one of its subdomains, or the README of github.com/openai/openai-guardrails-python.
-
Tell us where it is
We read it once now and again every week. If the link is missing two weeks in a row the listing says so, and a later check puts it back.
Agents send the same to POST /api/v1/verify as {"slug": "openai-guardrails", "url": "…"}, or call the verify_listing tool at /mcp. Ten checks an hour from one address. What we check. To announce the listing, get sharing assets for social media.


