arXiv:2501.15430cs.CL2025-01被引 1

针对罗伯特模型在识别非裔英语文本时的偏见问题,测试简单去偏技术效果。

Evaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models

  • 在罗伯特模型中应用去偏技术,缓解非裔英语文本误判
  • 正确处理数据集构建方法可显著降低方言群体间误判差异
  • 适合关注模型公平性与社会影响的研究者阅读

仇恨言论检测任务因训练数据中标注偏差,常对非裔美国人英语(AAE)方言文本产生偏见,导致正常AAE文本更易被误判为攻击性内容。本文评估了以往提出的简单去偏技术在基于RoBERTa的编码器中的表现。实验表明,这些技术的效果高度依赖于训练数据集的构建方式;但在充分考虑表征偏差的前提下,能有效降低不同方言子群体间的识别差异。研究强调了数据构建策略在实现模型公平性中的关键作用。

原文摘要 · Abstract (English)

The hate speech detection task is known to suffer from bias against African American English (AAE) dialect text, due to the annotation bias present in the underlying hate speech datasets used to train these models. This leads to a disparity where normal AAE text is more likely to be misclassified as abusive/hateful compared to non-AAE text. Simple debiasing techniques have been developed in the past to counter this sort of disparity, and in this work, we apply and evaluate these techniques in the scope of RoBERTa-based encoders. Experimental results suggest that the success of these techniques depends heavily on the methods used for training dataset construction, but with proper consideration of representation bias, they can reduce the disparity seen among dialect subgroups on the hate speech detection task.

仇恨言论检测去偏RoBERTa

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。