Skip to content

Guide

Use AI to moderate content

Moderation & safety · 4 min read

AI moderation uses an AI model (OpenAI Moderation) to check posts and, if you choose, replies when they're created or edited. It scores content in categories such as harassment, hate, and violence, and can hide or flag content that crosses your thresholds. It works alongside your keyword and spam filter rules, not instead of them.

Who can do this: Owners and admins turn AI moderation on and configure it. Moderators can view AI Activity and review held content. · Where: Web, in the admin console. · Plan: Available on every platform plan. AI moderation doesn't use your plan's AI allowance.

Turn on AI moderation

  1. Go to Admin → Settings → Content Safety. The page is titled Content Moderation.
  2. Under AI Moderation, turn on Enable AI moderation.
  3. Click Save Changes at the bottom of the page.

When the switch is off, no posts are sent to OpenAI. Content is sent to a third-party service, so review your privacy policy before you turn it on. See AI data privacy and limitations.

Configure AI Classification

The detailed settings live under Admin → Moderation → Automated Moderation, in the AI Classification section. You can also click Configure in Automated Moderation on the Content Safety page. Only owners and admins see this section.

Click Configure and set:

  • Mode: Shadow or Enforce (see the next section). New communities start in Shadow.
  • Scan: Top-level posts, Replies and comments, or both. Pick at least one. Replies generate 5–20× the moderation traffic of top-level posts.
  • Thresholds: when content is blocked or flagged, per category.

Click Save changes. Saving here doesn't change the on/off switch in Content Safety.

Shadow mode and Enforce mode

Mode What happens
Shadow The AI only records scores. Nothing is hidden and nothing is queued. Use it to see how the AI judges your community's content.
Enforce Content over a Block threshold is hidden immediately and sent for review. Content over a Flag threshold stays visible and waits for a moderator's decision.

When you switch to Enforce, you must tick a confirmation that flagged posts will be withheld from now on. Review shadow-mode results on AI Activity first.

Set AI moderation thresholds

Each category has a threshold from 0 to 1. A category triggers when the content's score is at or above the threshold, so a lower threshold is stricter.

  • Block (reject content): severe categories, such as Sexual content involving minors, Graphic violence, Self-harm instructions, Threatening hate speech, and Violent wrongdoing.
  • Flag (send to review queue): broader categories: Hate, Harassment, Sexual content, Violence, Self-harm, and Illicit behavior.

Tick a category to use it and set its value, or click Recommended defaults. Each tier is saved as a whole: categories you leave unticked in a tier no longer trigger. Advanced (JSON) shows the same settings as JSON.

Review content held by AI moderation

In Enforce mode, blocked and flagged content waits in Admin → Moderation → Pending Content, where Flagged for shows AI. Approve it to publish it or reject it to remove it. Blocked content stays hidden until you decide; flagged content stays visible meanwhile. For how to review it and what authors see, see Review held and pending content.

Check AI decisions on the AI Activity page

Admin → Moderation → AI Activity lists every AI moderation decision. It's a log, not a work queue: acting here doesn't change the content.

  • Each row shows the Content, Author, Categories, Score, Mode, Result, and Date. Results are Would block or Would flag (shadow, recorded only), Blocked, Flagged, Not enforced, or Passed.
  • Filter by mode, result, category, content type, and date range.
  • Click a row to see the model, thresholds, and Category scores.
  • Use Correct call or Wrong call (overturn) to rate the AI's decisions. Overturn rate shows how often moderators disagree.
  • By category shows how often each category fires and is overturned. A high overturn rate means that threshold is too loose. Click Adjust thresholds to change it.

Things to know about AI moderation

  • AI moderation checks the title and body text of posts and replies, when they're created and when they're edited. It doesn't check images, attachments, or videos.
  • Posts and replies by community admins and space moderators aren't sent to AI moderation.
  • Keyword, regex, and spam filter rules are set separately. See Set up auto-moderation rules.
  • AI moderation doesn't use your plan's AI allowance. See Understand AI usage and credits.

Related articles

Was this guide helpful?

Back to guides

Can't find what you need?

If you're a member of a community, its admins are the right people to ask. If you're building or running one on Mateflow, our team can help.

Start free trial