arXiv:2605.17187cs.CLcs.AI2026-05ACL

构建多语言社区违规检测基准,揭示大模型在多元规则下表现有限。

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

论文配图:PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
图 1 · 摘自论文原文
  • 将社区违规识别建模为多选任务,模拟真实人工审核流程。
  • 覆盖1989个Reddit社区、2885条规则,含13371个违规样本,支持9种语言。
  • 大模型表现仅略优于基线,通用规则更易识别,复杂情境仍具挑战。

社交媒体正向多元共治模式演进——由社区自主定义规范,不同社区对同一行为的容忍度各异。人工智能能否协助此类多元社区的治理?我们首次将该任务形式化为多选题,模拟真实人工审核:给定评论及其上下文,判断其违反了哪一条具体规则。为此,我们构建了PluRule,一个跨模态、多语言的基准数据集,涵盖1,989个Reddit社区、2,885条规则,共13,371个违规样本,覆盖9种语言。实验表明,当前最先进的视觉-语言模型表现不佳:即使使用性能优异的GPT-5.2模型,在高推理能力下也仅略优于简单基线。此外,模型规模扩大或上下文增加带来的提升极为有限。通用规则如礼貌性与自推广行为相对容易识别。结果表明,多元社区内容治理是语言模型面临的一项根本性挑战。相关代码与数据集已公开。

原文摘要 · Abstract (English)

Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community may be perfectly acceptable in another. Can AI models help moderate such pluralistic communities? We formalize the task as a multiple-choice problem, mirroring how human moderators operate in the real world: given a comment and its surrounding context, identify which specific rule, if any, is violated. We introduce PluRule, a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities spanning 2,885 rules in 9 languages. Using this benchmark, we show that state-of-the-art vision-language models struggle significantly: even GPT-5.2 with high reasoning performs only slightly better than a trivial baseline. We also find that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect. Our results show that moderation of pluralistic communities on social media is a fundamental challenge for language models. Our code and benchmark are publicly available.

内容审核多语言社区治理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。