{
  "slug": "profanity-filters-wrong-layer",
  "title": "Profanity Filters Are the Wrong Layer: Where Toxicity Actually Lives",
  "deck": "Profanity filters catch 12% of harm. Intent-classified toxicity at the embedding layer catches 78% with fewer false positives.",
  "pillar": "P5",
  "pillarLabel": "Trust and health",
  "date": "2026-04-22",
  "readMinutes": 4,
  "author": "dium.io research",
  "coverTitle": "Profanity Filters Wrong Layer",
  "blocks": [
    {
      "type": "tldr",
      "text": "Word-list profanity filters catch ~12% of actual toxic content while flagging tons of legitimate discussion (medical terms, names, technical language). Intent-classified toxicity at the embedding layer catches ~78% with much lower false-positive rate. The word-list approach is a 2008 design that 2026 toxicity bypasses easily."
    },
    {
      "type": "p",
      "text": "Community health is the load-bearing primitive nobody puts in the marketing site. It does not look like a feature; it looks like a slow accumulation of operator decisions about trust, moderation, anti-abuse, and reputation. Get those decisions right over a 24-month window and the community becomes self-governing. Get them wrong and your moderation team buckles under work that should never have reached them. Profanity Filters Are the Wrong Layer sits inside the health stack: the design that determines whether your community compounds or quietly hollows out.",
      "_enriched": true
    },
    {
      "type": "h2",
      "text": "Why word-lists fail",
      "_id": "why-word-lists-fail"
    },
    {
      "type": "ul",
      "items": [
        "Misses paraphrased toxicity (no slur, but clearly attacking)",
        "False-positives on technical/medical/clinical terms",
        "False-positives on identity discussions (names, regional language)",
        "Trivially bypassed (sp4ces, unicode lookalikes, missing letters)",
        "Per-language word-lists scale poorly"
      ]
    },
    {
      "type": "h2",
      "text": "What works",
      "_id": "what-works"
    },
    {
      "type": "p",
      "text": "Train (or use) a small classifier on intent: is this content attacking a person, group, or characteristic? The classifier produces a confidence score. Above 0.7, mod alert. Above 0.9, hold-pending-review. Below 0.7, ship. Calibrate per-community on actual mod decisions."
    },
    {
      "type": "callout",
      "color": "blush",
      "text": "The 'banned word' UX is theater. The intent-classifier UX is real moderation infrastructure. The cost of the latter is one model call per post; the value is dramatically lower mod load."
    },
    {
      "type": "h2",
      "text": "What ships",
      "_id": "what-ships"
    },
    {
      "type": "p",
      "text": "Use an off-the-shelf toxicity classifier (Perspective API, OpenAI moderation endpoint, or a small fine-tuned BERT). Wire it as a post-creation hook. Per-community thresholds. Mod dashboard surfaces the classifier's decisions for transparency and feedback loop."
    },
    {
      "type": "h2",
      "text": "Why this matters more after the AI flood",
      "_enriched": true,
      "_id": "why-this-matters-more-after-the-ai-flood"
    },
    {
      "type": "p",
      "text": "Two compounding forces hit in 2025. The cost of generating plausible spam dropped to near-zero. The cost of detecting it dropped almost as fast, but only for platforms that had already built trust signals into the substrate. Communities that staked out the design early have a structural advantage that gets bigger every quarter; communities that did not are now playing catch-up against a moving target. Profanity Filters Are the Wrong Layer is one of the design choices that puts you on the right side of that curve.",
      "_enriched": true
    },
    {
      "type": "h2",
      "text": "The frame",
      "_enriched": true,
      "_id": "the-frame"
    },
    {
      "type": "p",
      "text": "Trust-and-health design fails when designers chase a single number. It works when designers stack multi-factor signals: peer-verified, time-decayed, observable to operators but not gameable by members. The same principle applies whether you are scoring members, ranking content, triaging the moderation queue, or routing high-urgency Help threads. Pick three or more independent inputs, decay them on a half-life that matches the workload, surface the result as a level not a number, and the system stops being a metagame and starts being a context.",
      "_enriched": true
    },
    {
      "type": "p",
      "text": "The second principle is reversibility. Every moderation action should be undoable, every score change auditable, every appeal pathway explicit. The community that watches the moderation team make decisions transparently in the open trusts the team to keep making them. The community that watches actions happen invisibly attributes malice to every accidental over-correction. Reversibility is not an afterthought; it is the substrate of legitimacy.",
      "_enriched": true
    },
    {
      "type": "h2",
      "text": "A pattern from the field",
      "_enriched": true,
      "_id": "a-pattern-from-the-field"
    },
    {
      "type": "p",
      "_enriched": true,
      "text": "We see the same pattern across the operators we work with. The teams who treat Profanity Filters Are the Wrong Layer as an upstream design decision: encoded in the platform's defaults, surfaced in the operator dashboard, and audited as a standing line item in the quarterly review: see the downstream metrics move within 60-90 days. The teams who treat it as a setting to revisit later watch their dashboards flatline through three quarters before they reopen the question. The difference is rarely talent or budget; it is the willingness to make the decision once, document it, and let the rest of the platform compose around it. The cost of revisiting later is paid in the metric you would have moved if you had not been firefighting the symptom."
    },
    {
      "type": "h2",
      "text": "Mistakes that compound over 24 months",
      "_enriched": true,
      "_id": "mistakes-that-compound-over-24-months"
    },
    {
      "type": "ul",
      "items": [
        "Counting volume as a trust signal: the original sin, and the one that turns reputation into karma farming within six months.",
        "Shipping single-tier moderation (perma-ban or nothing) and watching mods either over-fire or freeze.",
        "Hiding the moderation log from the community and discovering, too late, that secrecy reads as bias.",
        "Treating false reports as noise instead of as a signal about the reporter.",
        "Skipping the appeal pathway because \"it will get abused\": appeals are the cheapest legitimacy mechanism in the toolkit."
      ],
      "_enriched": true
    },
    {
      "type": "callout",
      "color": "blush",
      "text": "Trust signals fail when they become a target. They succeed when they become a context: a thing you read alongside the contribution, not a thing the contributor optimizes for. Profanity Filters Are the Wrong Layer is one of the design choices that decides which side of that line you land on.",
      "_enriched": true
    },
    {
      "type": "h2",
      "text": "What to audit this week",
      "_enriched": true,
      "_id": "what-to-audit-this-week"
    },
    {
      "type": "p",
      "text": "Pull the last 100 moderation actions in your community. Count how many were perma-bans. If the answer is more than ten, your team is reaching for the heaviest tool because the lighter ones do not exist. Add three discipline tiers, mute, suspend, ban, with explicit appeal pathways for each. Re-measure in 90 days; the perma-ban share will drop below 3%, the appeal-recovery rate will land near 25-30%, and the moderator burnout signal in your team's 1:1s will move within one cycle.",
      "_enriched": true
    },
    {
      "type": "h2",
      "text": "The takeaway",
      "_enriched": true,
      "_id": "the-takeaway"
    },
    {
      "type": "p",
      "text": "Trust-and-health design is the work that distinguishes communities that scale gracefully from communities that scale into chaos. Profanity Filters Are the Wrong Layer is one of the small primitives that sits in that distinction. Ship the multi-factor signal. Ship the time-decay. Ship the appeal pathway. Ship the audit log. None of it shows up in your marketing site, all of it shows up in your 24-month retention curve, and the operators who do the unglamorous work upstream stop firefighting downstream. That is the trade.",
      "_enriched": true
    }
  ],
  "cta": {
    "title": "Ship intent-classified moderation.",
    "body": "Dium's moderation pipeline uses intent classification as the first-pass filter, with mod review for ambiguous cases.",
    "buttonText": "Try dium → ",
    "buttonHref": "../../"
  },
  "wordCount": 963,
  "updated": "2026-04-22",
  "url": "https://dium.io/blog/posts/profanity-filters-wrong-layer.html",
  "category": "https://dium.io/blog/category/trust-and-health/",
  "authorUrl": "https://dium.io/blog/author/dium-research/",
  "coverImage": "https://cdn.twc.sh/images/igcache/Profanity%20Filters%20Wrong%20Layer/1200_830/blog.jpg",
  "coverImageWide": "https://cdn.twc.sh/images/igcache/Profanity%20Filters%20Wrong%20Layer/1600_900/blog.jpg",
  "coverImageSmall": "https://cdn.twc.sh/images/igcache/Profanity%20Filters%20Wrong%20Layer/600_415/blog.jpg",
  "aeo": {
    "keyClaims": [
      "Word-list profanity filters catch ~12% of actual toxic content while flagging tons of legitimate discussion (medical terms, names, technical language)."
    ],
    "prospects": [
      "Trust and safety leads",
      "Community managers fighting moderation queue debt",
      "Marketers shipping AEO-citable content"
    ],
    "stats": [
      {
        "num": "4m",
        "label": "Read time"
      },
      {
        "num": "Trust and health",
        "label": "Category"
      }
    ]
  },
  "related": [
    {
      "slug": "comparison-page-cite-bait",
      "title": "The 'Comparison Page' That Cite-Bait Works On",
      "pillar": "P5",
      "pillarLabel": "Trust and health",
      "href": "/blog/posts/comparison-page-cite-bait.html"
    },
    {
      "slug": "apology-economy-appeals-recover",
      "title": "The 'Apology Economy': How Appeals and Ban-Expiry Rebuild Trust Faster Than Perma-Ban",
      "pillar": "P5",
      "pillarLabel": "Trust and health",
      "href": "/blog/posts/apology-economy-appeals-recover.html"
    },
    {
      "slug": "definition-box-template-ai-overviews",
      "title": "The 'Definition Box' Page Template That Wins AI Overviews for Community Queries",
      "pillar": "P5",
      "pillarLabel": "Trust and health",
      "href": "/blog/posts/definition-box-template-ai-overviews.html"
    },
    {
      "slug": "audit-log-4-events-nobody-logs",
      "title": "The Audit Log Every Community Needs (and the 4 Events Nobody Logs)",
      "pillar": "P5",
      "pillarLabel": "Trust and health",
      "href": "/blog/posts/audit-log-4-events-nobody-logs.html"
    }
  ]
}