arXiv:2603.22985cs.CLcs.CY2026-03被引 1

区分粗鲁语气与歧视内容,提升多模态内容审核准确性

Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation

  • 将仇恨言论细分为粗鲁语气和歧视内容两个维度
  • 联合使用粗粒度与细粒度标签,使模型误检率下降30%以上
  • 适合需要精准识别网络暴力的平台与研究者

当前多模态毒性评测通常采用单一二分类仇恨标签,这种粗略方法混淆了表达的两个本质特征:语气与内容。基于传播学理论,我们提出一种细粒度标注方案,区分两个可分离维度:不文明(粗鲁或轻蔑语气)和不容忍(攻击多元性并针对群体或身份的内容)。该方案应用于来自 Hateful Memes 数据集的 2,030 张表情包。我们在粗标签训练、跨标注体系迁移学习以及联合学习(结合粗粒度仇恨标签与细粒度标注)三种策略下评估多个视觉语言模型。结果表明,细粒度标注能补充现有粗标签,联合使用时显著提升整体模型性能。此外,采用细粒度方案训练的模型误差分布更均衡,对有害内容的漏检率明显降低(如 LLaVA-1.6-Mistral-7B 的 FNR-FPR 从 0.74 降至 0.42;Qwen2.5-VL-7B 从 0.54 降至 0.28)。本工作通过提升数据质量,推动以数据为中心的内容审核方法,为构建更可靠、准确的多模态审核系统提供可行路径。

原文摘要 · Abstract (English)

Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: tone and content. Drawing on communication science theory, we introduce a fine-grained annotation scheme that distinguishes two separable dimensions: incivility (rude or dismissive tone) and intolerance (content that attacks pluralism and targets groups or identities) and apply it to 2,030 memes from the Hateful Memes dataset. We evaluate different vision-language models under coarse-label training, transfer learning across label schemes and a joint learning approach that combines the coarse hatefulness label with our fine-grained annotations. Our results show that fine-grained annotations complement existing coarse labels and, when used jointly, improve overall model performance. Moreover, models trained with the fine-grained scheme exhibit more balanced moderation-relevant error profiles and are less prone to under-detection of harmful content than models trained on hatefulness labels alone (FNR-FPR, the difference between false negative and false positive rates: 0.74 to 0.42 for LLaVA-1.6-Mistral-7B; 0.54 to 0.28 for Qwen2.5-VL-7B). This work contributes to data-centric approaches in content moderation by improving the reliability and accuracy of moderation systems through enhanced data quality. Overall, combining both coarse and fine-grained labels provides a practical route to more reliable multimodal moderation.

内容审核多模态细粒度标注仇恨言论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。