arXiv:2410.10332cs.CLcs.AI2024-10NAACL被引 2

发现仇恨言论检测模型对不同目标群体存在系统性偏见。

Disentangling Hate Across Target Identities

  • 基于目标身份提及自动提升仇恨评分,与内容无关。
  • 模型混淆仇恨度与情绪极性,误判愤怒表达为仇恨。
  • 刻板印象强度越高,预测准确率越差,风险加剧。

仇恨言论(HS)分类器在检测针对不同目标身份的仇恨表达时表现不一,且预测的仇恨程度评分存在系统性偏差。基于两个近期提出的功能测试数据集,我们量化分析了多种因素对HS预测的影响。在主流工业和学术模型上的实验表明,仅因提及特定目标身份,模型便会赋予更高的仇恨评分。此外,模型常将仇恨度与情绪极性混淆,导致对表达愤怒或反对仇恨言论的帖子误判为仇恨内容。该结果令人担忧,因为构建这些检测器可能反而伤害我们本欲保护的弱势群体。受社会心理学理论启发的研究进一步揭示,仇恨程度预测的准确性与刻板印象强度高度相关。

原文摘要 · Abstract (English)

Hate speech (HS) classifiers do not perform equally well in detecting hateful expressions towards different target identities. They also demonstrate systematic biases in predicted hatefulness scores. Tapping on two recently proposed functionality test datasets for HS detection, we quantitatively analyze the impact of different factors on HS prediction. Experiments on popular industrial and academic models demonstrate that HS detectors assign a higher hatefulness score merely based on the mention of specific target identities. Besides, models often confuse hatefulness and the polarity of emotions. This result is worrisome as the effort to build HS detectors might harm the vulnerable identity groups we wish to protect: posts expressing anger or disapproval of hate expressions might be flagged as hateful themselves. We also carry out a study inspired by social psychology theory, which reveals that the accuracy of hatefulness prediction correlates strongly with the intensity of the stereotype.

仇恨言论模型偏见社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。