从历史记录中提取Reddit社区隐性内容审核标准
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
- 用可解释架构分析历史审核数据,生成词汇评分表
- 提取的规则能复现神经模型效果,准确率相当
- 揭示社区间语言容忍度、话题限制等隐藏差异
有效的内容审核系统需要明确的分类标准,但像子版块这样的在线社区往往采用多样且隐性的规范。本文提出一种新方法,通过可解释架构从历史审核数据中识别并提取这些隐性标准。我们将审核标准表示为与内容删除相关的词汇表达评分表,实现不同社区间的系统性对比。实验表明,所提取的词汇模式能够有效复现神经审核模型的性能,同时提供透明的决策过程洞察。结果矩阵揭示了看似共享的规范在实际执行中的显著差异,发现了此前未记录的审核模式,包括社区特有的语言容忍度、主题限制特征以及毒性言论分类的潜在子类别。
原文摘要 · Abstract (English)
Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work introduces a novel approach to identify and extract these implicit criteria from historical moderation data using an interpretable architecture. We represent moderation criteria as score tables of lexical expressions associated with content removal, enabling systematic comparison across different communities. Our experiments demonstrate that these extracted lexical patterns effectively replicate the performance of neural moderation models while providing transparent insights into decision-making processes. The resulting criteria matrix reveals significant variations in how seemingly shared norms are actually enforced, uncovering previously undocumented moderation patterns including community-specific tolerances for language, features for topical restrictions, and underlying subcategories of the toxic speech classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。