构建社区视角的毒性语言数据集,提升内容审核的包容性
ModelCitizens: Representing Community Voices in Online Safety
- 基于6.8K条社交媒体帖子与40K条标注,融合多元身份群体观点
- 现有工具在上下文增强后性能下降,新模型在分布内测试中领先5.5%
- 适合关注公平性、社区参与的内容安全研究者使用
自动识别网络暴力语言对营造安全包容的在线空间至关重要,但其判断高度主观,受社群规范与生活经验影响。现有模型通常将多样标注合并为单一真实标签,忽略了如重获语言等情境化毒性概念。为此,我们提出MODELCITIZENS,包含6.8K条社交媒体帖子与40K条跨身份群体的毒性标注,并通过大模型生成对话场景以增强语境。主流检测工具(如OpenAI Moderation API、GPT-o4-mini)在该数据集上表现不佳,且在上下文增强后进一步下降。我们还发布了基于LLaMA和Gemma的LLAMACITIZEN-8B与GEMMACITIZEN-12B模型,在分布内评估中比GPT-o4-mini高出5.5%。结果强调了社区导向标注与建模在包容性内容管理中的重要性。数据、模型与代码已开源。
原文摘要 · Abstract (English)
Automatic toxic language detection is critical for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and lived experience. Existing toxicity detection models are typically trained on annotations that collapse diverse annotator perspectives into a single ground truth, erasing important context-specific notions of toxicity such as reclaimed language. To address this, we introduce MODELCITIZENS, a dataset of 6.8K social media posts and 40K toxicity annotations across diverse identity groups. To capture the role of conversational context on toxicity, typical of social media posts, we augment MODELCITIZENS posts with LLM-generated conversational scenarios. State-of-the-art toxicity detection tools (e.g. OpenAI Moderation API, GPT-o4-mini) underperform on MODELCITIZENS, with further degradation on context-augmented posts. Finally, we release LLAMACITIZEN-8B and GEMMACITIZEN-12B, LLaMA- and Gemma-based models finetuned on MODELCITIZENS, which outperform GPT-o4-mini by 5.5% on in-distribution evaluations. Our findings highlight the importance of community-informed annotation and modeling for inclusive content moderation. The data, models and code are available at https://github.com/asuvarna31/modelcitizens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。