用解释方法检测算法偏见,揭示公平性问题。
Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration
- 将局部后验解释方法整合成流水线,用于发现公平性问题
- 发现解释方法在群体公平评估中效果不一致,需谨慎使用
- 适合关注算法公平性与可解释性的研究人员
随着人工智能在影响人类生活的领域日益普及,公平性与透明性问题愈发受到关注,尤其涉及受保护群体时。近年来,可解释性与公平性的交叉成为推动负责任AI系统的重要方向。本文探讨如何利用可解释性方法检测并解读不公平现象。提出一个集成局部后验解释方法的分析流水线,以获取与公平性相关的洞见。在设计过程中,识别并解决了若干关键问题:分配公平与程序公平的关系、移除受保护属性的影响、不同解释方法结果的一致性与质量、局部解释聚合策略对群体公平评估的影响,以及解释作为偏见检测器的整体可信度。实验结果表明,解释方法在公平性探索中具有潜力,但需审慎考虑上述关键因素。
原文摘要 · Abstract (English)
As Artificial Intelligence (AI) is increasingly used in areas that significantly impact human lives, concerns about fairness and transparency have grown, especially regarding their impact on protected groups. Recently, the intersection of explainability and fairness has emerged as an important area to promote responsible AI systems. This paper explores how explainability methods can be leveraged to detect and interpret unfairness. We propose a pipeline that integrates local post-hoc explanation methods to derive fairness-related insights. During the pipeline design, we identify and address critical questions arising from the use of explanations as bias detectors such as the relationship between distributive and procedural fairness, the effect of removing the protected attribute, the consistency and quality of results across different explanation methods, the impact of various aggregation strategies of local explanations on group fairness evaluations, and the overall trustworthiness of explanations as bias detectors. Our results show the potential of explanation methods used for fairness while highlighting the need to carefully consider the aforementioned critical aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。