arXiv:2505.14483cs.CL2025-05EMNLP被引 19

用专家协作框架实现跨社区可解释的内容审核,无需每社区单独训练。

MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance

  • 设计模块化专家系统,动态分配内容审核任务
  • 在30个未见子版块上达到0.72和0.67的Micro-F1分数
  • 输出简洁可靠的解释,适合需要透明治理的平台

大型语言模型在识别在线社区中的有害内容方面展现出巨大潜力。然而,现有的审核方法需为每个社区单独部署模型,且决策过程不透明,限制了实际应用。本文提出混合审核专家框架(MoMoE),一种模块化、跨社区的内容审核机制,可提供事后解释并支持规模化应用。MoMoE 调度四种操作:分配、预测、聚合与解释,并具体实现七个社区专用专家(MoMoE-Community)和五个规范违规专家(MoMoE-NormVio)。在30个未见过的子版块上,最佳变体分别获得0.72和0.67的Micro-F1得分,性能匹配或超越强基线微调模型,同时持续生成简明可靠的解释。尽管社区专用专家达到最高峰值准确率,但规范违规专家在跨域场景中表现更稳定。结果表明,MoMoE 可实现无需每社区微调的可扩展、透明审核。更广泛地,轻量级可解释专家集成或可引导未来NLP与人机交互在可信人机协同在线治理方面的研究。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown great potential in flagging harmful content in online communities. Yet, existing approaches for moderation require a separate model for every community and are opaque in their decision-making, limiting real-world adoption. We introduce Mixture of Moderation Experts (MoMoE), a modular, cross-community framework that adds post-hoc explanations to scalable content moderation. MoMoE orchestrates four operators -- Allocate, Predict, Aggregate, Explain -- and is instantiated as seven community-specialized experts (MoMoE-Community) and five norm-violation experts (MoMoE-NormVio). On 30 unseen subreddits, the best variants obtain Micro-F1 scores of 0.72 and 0.67, respectively, matching or surpassing strong fine-tuned baselines while consistently producing concise and reliable explanations. Although community-specialized experts deliver the highest peak accuracy, norm-violation experts provide steadier performance across domains. These findings show that MoMoE yields scalable, transparent moderation without needing per-community fine-tuning. More broadly, they suggest that lightweight, explainable expert ensembles can guide future NLP and HCI research on trustworthy human-AI governance of online communities.

内容审核专家系统可解释AI人机治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。