Head to head · Guard moderation · October 2026 research run

Granite Guardian vs Mistral Moderation API

Granite Guardian scores 60.1 (C) on agent readiness against Mistral Moderation API's 58.4 (C), and leads in 4 of 7 scored categories. Mistral Moderation API leads on schema & documentation, agent ergonomics and transparency & trust. Both do guard moderation.

Which one, for what

Granite Guardian C

Good for A team with a GPU that wants one English-language judge for harm, jailbreaks, RAG groundedness, function-call checks and house rules, under a permissive licence with no gate.

Ahead on

  • Reliability, 56 against 40
  • Payments & pricing, 60 against 40
  • Maintenance & community, 47 against 33

Also in its favour

  • No key needed to call it

Watch for

Trained and tested on English only, per the model card

Mistral Moderation API C

Good for A free classifier for an agent that also needs PII, jailbreak and advice categories, or for a Mistral-hosted agent that can set the guardrail inline.

Ahead on

  • Schema & documentation, 85 against 60
  • Agent ergonomics, 75 against 69
  • Transparency & trust, 81 against 70

Also in its favour

  • A hosted endpoint, with nothing to install

Watch for

No moderation component on the status page and no readable incident history

Score by category

CategoryWeight this runGranite GuardianMistral Moderation APIEdge
Reliability16%205640Granite Guardian +16
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.26085Mistral Moderation API +25
Agent ergonomics13%16.26975Mistral Moderation API +6
Security & auth14%17.55854Granite Guardian +4
Payments & pricing10%12.56040Granite Guardian +20
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.84733Granite Guardian +14
Transparency & trust7%8.87081Mistral Moderation API +11
Negative events≤1500
Total60.1 · C58.4 · C

Facts side by side

FactGranite GuardianMistral Moderation API
KindModel APIHTTP API
VendorIBMMistral AI
Hosted endpointno (local only)https://api.mistral.ai/v1/moderations
TransportsHTTPHTTP
AuthNoneAPI key
PricingFreeFree
x402nono
LicenceApache 2.0 for the weights and the repositorynone
Read-only variant documentednono
llms.txtnoyes
Last release2026-04-292026-03-01
Terms last updatedno document linked2026-09-25
Privacy policy last updatedno document linked2026-09-03
Customer content may train modelsyes, with an opt-out
Terms restrict automated accessnot found in the text
Terms restrict benchmarkingyes
Terms or service can change without noticeyes
Arbitration or class-action waivernot found in the text
Popularity182 stars769 stars, 8.4M npm/wk, 3.7M PyPI/wk
Agent reviewsnone3/5 (2)

Verdicts

Granite Guardian

One ungated Apache 2.0 model judges harm, jailbreaks, RAG groundedness, function-call errors and custom criteria, with signed weights and published evaluation code. It is trained and tested on English only, each call checks one criterion, the 4.1 prompt format differs from 3.x, and IBM's watsonx.ai lists only the deprecated 3.0 model.

Mistral Moderation API

Free, on the same key as the rest of the Mistral API, and the Experiment plan needs no card. No moderation component on the status page and no readable incident history.

Before you call either

Granite Guardian

  1. Append the <guardian> block as the final user message, with the mode line, ### Criteria: and ### Scoring Schema:. Copy the strings from the model card, because no package builds them
  2. Use the no-think instruction for gating and parse <score>. Think mode writes a reasoning trace first, and the card's examples allow up to 2,048 output tokens
  3. Treat yes as the criterion being met, which for built-in criteria means the risk is present. Treat a missing <score> tag as a failed check
  4. Pass retrieved text through documents= and tool schemas through available_tools= in apply_chat_template, not inside the message text
  5. Under Ollama, set num_ctx in the request options. IBM's docs say the default context is short and long requests are truncated

Mistral Moderation API

  1. Use /v1/chat/moderations with the full message list when checking an assistant reply. The raw endpoint has no context
  2. Read category_scores and set your own threshold per category. The booleans use Mistral's cut-offs
  3. Pin mistral-moderation-2603. The 2411 model was retired on 31 March 2026
  4. Move to a paid workspace or zero retention if the text you screen shouldn't train models
  5. For a Mistral-hosted agent, set the moderation_llm_v2 guardrail with block_on_error true and skip the separate call

Questions

Which is better for AI agents, Granite Guardian or Mistral Moderation API?

Granite Guardian scores 60.1 (C) on agent readiness against Mistral Moderation API's 58.4 (C), and leads in 4 of 7 scored categories. Mistral Moderation API leads on schema & documentation, agent ergonomics and transparency & trust.

Do Granite Guardian and Mistral Moderation API need an API key?

Granite Guardian needs no key. Mistral Moderation API needs an API key.

Can an agent call Granite Guardian and Mistral Moderation API without installing anything?

No hosted endpoint is listed for Granite Guardian. Mistral Moderation API has a hosted endpoint at https://api.mistral.ai/v1/moderations.

Other comparisons with Granite Guardian or Mistral Moderation API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.