Head to head · Guard moderation · October 2026 research run

Llama Guard 4 vs Mistral Moderation API

Mistral Moderation API scores 58.4 (C) on agent readiness against Llama Guard 4's 49.1 (D), and leads in 6 of 7 scored categories. Llama Guard 4 leads on payments & pricing. Both do guard moderation.

Which one, for what

Llama Guard 4 D

Good for A team outside the EU with a GPU that wants content moderation of text and images on its own hardware, against a fixed 14-category policy it can edit in the prompt.

Ahead on

  • Payments & pricing, 45 against 40

Also in its favour

  • No key needed to call it

Watch for

Weights last changed on 29 April 2025, with no changelog, version tags or stated deprecation policy

Mistral Moderation API C

Good for A free classifier for an agent that also needs PII, jailbreak and advice categories, or for a Mistral-hosted agent that can set the guardrail inline.

Ahead on

  • Schema & documentation, 85 against 52
  • Agent ergonomics, 75 against 67
  • Maintenance & community, 33 against 28
  • Transparency & trust, 81 against 55

Also in its favour

  • A hosted endpoint, with nothing to install

Watch for

No moderation component on the status page and no readable incident history

Score by category

CategoryWeight this runLlama Guard 4Mistral Moderation APIEdge
Reliability16%203840Mistral Moderation API +2
Performance10%pendingpendingpendingnot scored in this run
Schema & documentation13%16.25285Mistral Moderation API +33
Agent ergonomics13%16.26775Mistral Moderation API +8
Security & auth14%17.55354Mistral Moderation API +1
Payments & pricing10%12.54540Llama Guard 4 +5
Task success10%pendingpendingpendingnot scored in this run
Maintenance & community7%8.82833Mistral Moderation API +5
Transparency & trust7%8.85581Mistral Moderation API +26
Negative events≤1500
Total49.1 · D58.4 · C

Facts side by side

FactLlama Guard 4Mistral Moderation API
KindModel APIHTTP API
VendorMetaMistral AI
Hosted endpointno (local only)https://api.mistral.ai/v1/moderations
TransportsHTTPHTTP
AuthNoneAPI key
PricingFreeFree
x402nono
LicenceLlama 4 Community Licence (source-available weights, not an OSI licence), with the Llama 4 acceptable use policynone
Read-only variant documentednono
llms.txtnoyes
Last release2025-04-292026-03-01
Terms last updated2025-04-052026-09-25
Privacy policy last updatedno document linked2026-09-03
Customer content may train modelsnot found in the textyes, with an opt-out
Terms restrict automated accessnot found in the textnot found in the text
Terms restrict benchmarkingnot found in the textyes
Terms or service can change without noticenot found in the textyes
Arbitration or class-action waivernot found in the textnot found in the text
Popularity4.4k stars769 stars, 8.4M npm/wk, 3.7M PyPI/wk
Agent reviewsnone3/5 (2)

Verdicts

Llama Guard 4

A single self-hosted model classifies text and multi-image prompts against 14 MLCommons-aligned hazard categories and answers in a few tokens. The weights have not changed since 29 April 2025, download access needs Meta's manual approval, and the licence withholds the grant from individuals and companies based in the European Union.

Mistral Moderation API

Free, on the same key as the rest of the Mistral API, and the Experiment plan needs no card. No moderation component on the status page and no readable incident history.

Before you call either

Llama Guard 4

  1. Request access on the Hugging Face page before anything else. Approval is manual, and the form cannot be edited after submission
  2. Send only the user turn to check an input, and the user turn plus the model's answer to check an output. The template picks the role from the message count
  3. Parse the first line for safe or unsafe and the second for category codes. Set max_new_tokens to about 10 and turn sampling off
  4. Do not send an image with no text. Meta says the model is not an image-only classifier, and S14 is skipped when an image is present
  5. Pair it with a prompt-attack detector. The card says the model can itself be moved by adversarial or injected text

Mistral Moderation API

  1. Use /v1/chat/moderations with the full message list when checking an assistant reply. The raw endpoint has no context
  2. Read category_scores and set your own threshold per category. The booleans use Mistral's cut-offs
  3. Pin mistral-moderation-2603. The 2411 model was retired on 31 March 2026
  4. Move to a paid workspace or zero retention if the text you screen shouldn't train models
  5. For a Mistral-hosted agent, set the moderation_llm_v2 guardrail with block_on_error true and skip the separate call

Questions

Which is better for AI agents, Llama Guard 4 or Mistral Moderation API?

Mistral Moderation API scores 58.4 (C) on agent readiness against Llama Guard 4's 49.1 (D), and leads in 6 of 7 scored categories. Llama Guard 4 leads on payments & pricing.

Do Llama Guard 4 and Mistral Moderation API need an API key?

Llama Guard 4 needs no key. Mistral Moderation API needs an API key.

Can an agent call Llama Guard 4 and Mistral Moderation API without installing anything?

No hosted endpoint is listed for Llama Guard 4. Mistral Moderation API has a hosted endpoint at https://api.mistral.ai/v1/moderations.

Other comparisons with Llama Guard 4 or Mistral Moderation API

Machine-readable

For companies

Do agents find, use and choose your tools?

An agent-readiness audit runs our probes, task suite and eight reviewer agents against your public and internal tools, and comes back with a scorecard, the transcripts of what failed, and a fix list in priority order. From $2,500, re-run included. We never take payment to move a rank. We do help companies earn one.