arXiv:2505.01198cs.CLcs.AI2025-05中稿 · ACM Conference on …被引 3

发现解释方法对不同性别群体存在显著性能差异,可能加剧算法偏见。

Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods

  • 测试三种任务五种语言模型,检验后验特征归因法的公平性
  • 方法在忠实性、鲁棒性和复杂度上均对特定性别表现更差
  • 即使使用无偏数据训练,偏差仍存在,提示需纳入监管考量

尽管解释方法的应用与评估研究持续扩展,但其在不同子群体间性能差异的公平性问题却常被忽视。本文通过在三个任务和五种语言模型上验证,发现广泛使用的后验特征归因方法在忠实性、鲁棒性和复杂度方面存在显著性别差异。这些差异在模型经无偏数据预训练或微调后依然存在,表明其并非仅由训练数据偏见导致。结果凸显了在开发与应用可解释性方法时,必须关注解释公平性,否则可能在高风险场景中对特定群体造成系统性不利影响。此外,研究强调应将解释公平性纳入监管框架,与模型整体公平性和可解释性并列考虑。

原文摘要 · Abstract (English)

While research on applications and evaluations of explanation methods continues to expand, fairness of the explanation methods concerning disparities in their performance across subgroups remains an often overlooked aspect. In this paper, we address this gap by showing that, across three tasks and five language models, widely used post-hoc feature attribution methods exhibit significant gender disparity with respect to their faithfulness, robustness, and complexity. These disparities persist even when the models are pre-trained or fine-tuned on particularly unbiased datasets, indicating that the disparities we observe are not merely consequences of biased training data. Our results highlight the importance of addressing disparities in explanations when developing and applying explainability methods, as these can lead to biased outcomes against certain subgroups, with particularly critical implications in high-stakes contexts. Furthermore, our findings underscore the importance of incorporating the fairness of explanations, alongside overall model fairness and explainability, as a requirement in regulatory frameworks.

解释公平性性别偏见可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。