arXiv:2510.06350cs.CYcs.AI2025-10AAAI

用问答模型精准匹配评论与社区规则,提升内容审核透明度。

Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation

  • 基于问答框架,实时关联评论与社区规则
  • 在Reddit和Lemmy数据上超越现有方法,准确率显著提升
  • 支持新社区和新规则快速适配,适合动态治理场景

在线社区依赖平台政策与社区自定规则来界定可接受行为并维持秩序。然而,这些规则在不同社区间差异巨大,随时间演变且执行不一致,给透明度、治理和自动化带来挑战。本文提出ModQ,一种新型问答框架,用于规则敏感的内容审核。不同于以往的分类或生成方法,ModQ在推理时依赖完整的社区规则集,识别最适用的规则。我们实现两种模型变体——抽取式和多选式问答,并在Reddit和Lemmy的大规模数据集上训练,后者由公开的审核日志和规则描述构建。两种模型在识别违规行为方面均优于现有最优基线,同时保持轻量与可解释性。值得注意的是,ModQ模型对未见过的社区和规则具有良好泛化能力,适用于低资源环境与动态治理场景。

原文摘要 · Abstract (English)

Online communities rely on a mix of platform policies and community-authored rules to define acceptable behavior and maintain order. However, these rules vary widely across communities, evolve over time, and are enforced inconsistently, posing challenges for transparency, governance, and automation. In this paper, we model the relationship between rules and their enforcement at scale, introducing ModQ, a novel question-answering framework for rule-sensitive content moderation. Unlike prior classification or generation-based approaches, ModQ conditions on the full set of community rules at inference time and identifies which rule best applies to a given comment. We implement two model variants - extractive and multiple-choice QA - and train them on large-scale datasets from Reddit and Lemmy, the latter of which we construct from publicly available moderation logs and rule descriptions. Both models outperform state-of-the-art baselines in identifying moderation-relevant rule violations, while remaining lightweight and interpretable. Notably, ModQ models generalize effectively to unseen communities and rules, supporting low-resource moderation settings and dynamic governance environments.

内容审核问答模型规则匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。