arXiv:2509.22291cs.CLcs.AI2025-09被引 2

输入解释能检测仇恨言论中的偏见,但难选公平模型。

Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?

  • 用输入级解释定位模型偏见预测
  • 解释可辅助训练时降低偏见,但无法可靠选公平模型
  • 首次系统量化分析解释与公平性的关系

自然语言处理模型常在训练数据中复制或放大社会偏见,引发公平性担忧。同时其黑箱特性使用户难以识别偏见预测,开发者也难以有效缓解。尽管有研究认为输入解释有助于发现和减轻偏见,但其在确保公平性方面的可靠性仍存疑。现有公平NLP中的可解释性研究多为定性,缺乏大规模定量分析。本文首次系统研究仇恨言论检测中可解释性与公平性的关系,涵盖编码器-和解码器仅模型,考察三个维度:(1) 识别偏见预测,(2) 选择公平模型,(3) 训练期间缓解偏见。结果表明,输入解释可有效检测偏见预测,并作为训练阶段减少偏见的有用监督信号,但对在候选模型间选择公平模型不可靠。代码已开源。

原文摘要 · Abstract (English)

Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes it difficult for users to recognize biased predictions and for developers to effectively mitigate them. While some studies suggest that input-based explanations can help detect and mitigate bias, others question their reliability in ensuring fairness. Existing research on explainability in fair NLP has been predominantly qualitative, with limited large-scale quantitative analysis. In this work, we conduct the first systematic study of the relationship between explainability and fairness in hate speech detection, focusing on both encoder- and decoder-only models. We examine three key dimensions: (1) identifying biased predictions, (2) selecting fair models, and (3) mitigating bias during model training. Our findings show that input-based explanations can effectively detect biased predictions and serve as useful supervision for reducing bias during training, but they are unreliable for selecting fair models among candidates.Our code is available at https://github.com/Ewanwong/fairness_x_explainability.

可解释性公平性仇恨言论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。