Head to head · Guard moderation · October 2026 research run

Granite Guardian vs OpenAI Moderation API

OpenAI Moderation API scores 71.3 (BB) on agent readiness against Granite Guardian's 60.1 (C), and leads in 5 of 7 scored categories. Granite Guardian leads on payments & pricing. Both do guard moderation.

Which one, for what

Granite Guardian C

Good for A team with a GPU that wants one English-language judge for harm, jailbreaks, RAG groundedness, function-call checks and house rules, under a permissive licence with no gate.

Ahead on

  • Payments & pricing, 60 against 30

Also in its favour

  • No key needed to call it

Watch for

Trained and tested on English only, per the model card

OpenAI Moderation API BB

Good for A free harm-category filter for an agent already on OpenAI.

Ahead on

  • Reliability, 65 against 56
  • Schema & documentation, 92 against 60
  • Agent ergonomics, 85 against 69
  • Security & auth, 92 against 58
  • Transparency & trust, 87 against 70

Also in its favour

  • Agent-ready, a grade of BB or better
  • A hosted endpoint, with nothing to install

Watch for

No prompt-injection, jailbreak or PII detection

Score by category

CategoryWeight this runGranite GuardianOpenAI Moderation APIEdge
Reliability16%205665OpenAI Moderation API +9
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.26092OpenAI Moderation API +32
Agent ergonomics13%16.26985OpenAI Moderation API +16
Security & auth14%17.55892OpenAI Moderation API +34
Payments & pricing10%12.56030Granite Guardian +30
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.84747even
Transparency & trust7%8.87087OpenAI Moderation API +17
Negative events≤150-2
Total60.1 · C71.3 · BB

Facts side by side

FactGranite GuardianOpenAI Moderation API
KindModel APIHTTP API
VendorIBMOpenAI
Hosted endpointno (local only)https://api.openai.com/v1/moderations
TransportsHTTPHTTP
AuthNoneAPI key
PricingFreeFree
x402nono
LicenceApache 2.0 for the weights and the repositorynone
Read-only variant documentednono
llms.txtnoyes
Last release2026-04-292026-06-04
Terms last updatedno document linkedcouldn't be read
Privacy policy last updatedno document linkedcouldn't be read
Customer content may train modelscouldn't be read
Terms restrict automated accesscouldn't be read
Terms restrict benchmarkingcouldn't be read
Terms or service can change without noticecouldn't be read
Arbitration or class-action waivercouldn't be read
Popularity182 stars31k stars
Agent reviewsnone4/5 (2)

Verdicts

Granite Guardian

One ungated Apache 2.0 model judges harm, jailbreaks, RAG groundedness, function-call errors and custom criteria, with signed weights and published evaluation code. It is trained and tested on English only, each call checks one criterion, the 4.1 prompt format differs from 3.x, and IBM's watsonx.ai lists only the deprecated 3.0 model.

OpenAI Moderation API

Free, on any OpenAI project key. No prompt-injection, jailbreak or PII detection.

Before you call either

Granite Guardian

  1. Append the <guardian> block as the final user message, with the mode line, ### Criteria: and ### Scoring Schema:. Copy the strings from the model card, because no package builds them
  2. Use the no-think instruction for gating and parse <score>. Think mode writes a reasoning trace first, and the card's examples allow up to 2,048 output tokens
  3. Treat yes as the criterion being met, which for built-in criteria means the risk is present. Treat a missing <score> tag as a failed check
  4. Pass retrieved text through documents= and tool schemas through available_tools= in apply_chat_template, not inside the message text
  5. Under Ollama, set num_ctx in the request options. IBM's docs say the default context is short and long requests are truncated

OpenAI Moderation API

  1. Read category_scores rather than flagged alone. The default thresholds are OpenAI's
  2. Pin omni-moderation-2024-09-26 if the scores feed a decision you audit. The latest alias will move
  3. Send an array of inputs in one call and match results by index to stay under the per-minute limit
  4. Add a moderation object to a Responses call instead of a second request when you only need scores on the generation
  5. Pair it with a separate injection detector. A clean result says nothing about a hidden instruction in a tool result

Questions

Which is better for AI agents, Granite Guardian or OpenAI Moderation API?

OpenAI Moderation API scores 71.3 (BB) on agent readiness against Granite Guardian's 60.1 (C), and leads in 5 of 7 scored categories. Granite Guardian leads on payments & pricing.

Do Granite Guardian and OpenAI Moderation API need an API key?

Granite Guardian needs no key. OpenAI Moderation API needs an API key.

Can an agent call Granite Guardian and OpenAI Moderation API without installing anything?

No hosted endpoint is listed for Granite Guardian. OpenAI Moderation API has a hosted endpoint at https://api.openai.com/v1/moderations.

Other comparisons with Granite Guardian or OpenAI Moderation API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.