让每个人和群体都获得公平的可操作解释,提升模型决策透明度。
Fair Recourse for All: Ensuring Individual and Group Fairness in Counterfactual Explanations
- 用强化学习生成兼顾个体与群体公平的反事实解释
- 在三个数据集上实现公平性与解释质量的平衡
- 适合关注算法公平性与可解释性的研究者
可解释人工智能(XAI)对提升机器学习模型透明度至关重要。其中,反事实解释(CFs)因其能展示输入特征如何影响模型决策,从而为用户提供可操作的改进路径而尤为关键。确保具有相似属性的个体以及不同受保护群体(如人口统计学群体)获得相似且可行的解释,是实现可信、公平决策的核心。本文直接应对该挑战,提出同时保障个体公平、群体公平及二者的混合公平性。我们将问题建模为优化任务,提出一种模型无关的基于强化学习的方法,同时满足个体与群体层面的公平约束——这两类目标通常被视为相互独立。我们扩展了常用于审计机器学习模型的公平性指标,如“等效可操作性”和“跨个体与群体的等效有效性”。在三个基准数据集上的评估表明,该方法能有效实现个体与群体公平性,同时保持解释在接近性和合理性方面的高质量,并分别量化了不同层次公平性的代价。本工作为混合公平性及其在XAI中的作用与影响开启了更广泛的讨论。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence (XAI) is becoming increasingly essential for enhancing the transparency of machine learning (ML) models. Among the various XAI techniques, counterfactual explanations (CFs) hold a pivotal role due to their ability to illustrate how changes in input features can alter an ML model's decision, thereby offering actionable recourse to users. Ensuring that individuals with comparable attributes and those belonging to different protected groups (e.g., demographic) receive similar and actionable recourse options is essential for trustworthy and fair decision-making. In this work, we address this challenge directly by focusing on the generation of fair CFs. Specifically, we start by defining and formulating fairness at: 1) individual fairness, ensuring that similar individuals receive similar CFs, 2) group fairness, ensuring equitable CFs across different protected groups and 3) hybrid fairness, which accounts for both individual and broader group-level fairness. We formulate the problem as an optimization task and propose a novel model-agnostic, reinforcement learning based approach to generate CFs that satisfy fairness constraints at both the individual and group levels, two objectives that are usually treated as orthogonal. As fairness metrics, we extend existing metrics commonly used for auditing ML models, such as equal choice of recourse and equal effectiveness across individuals and groups. We evaluate our approach on three benchmark datasets, showing that it effectively ensures individual and group fairness while preserving the quality of the generated CFs in terms of proximity and plausibility, and quantify the cost of fairness in the different levels separately. Our work opens a broader discussion on hybrid fairness and its role and implications for XAI and beyond CFs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。