提出解释公平性统一框架,解决模型推理不公问题
Fairness of Explanations in Artificial Intelligence (AI): A Unifying Framework, Axioms, and Future Direction toward Responsible AI

- 构建条件不变性框架,定义解释公平的数学标准
- 揭示解释不公三大生成机制,如特征表示偏差
- 提供可操作的六步审计流程,适合合规与伦理审查
机器学习广泛应用于刑事司法、医疗、信贷和就业等高风险决策。学界发展出两个独立方向:算法公平性关注结果平等,可解释AI(XAI)关注推理可理解性。本文首次揭示二者交汇处的盲点——模型虽满足所有输出公平标准,但其推理过程仍存在严重程序性偏见。为此提出‘条件不变性框架’,形式化解释公平为:对所有相关输入,解释分布应与敏感属性无关。该框架统一了现有解释公平度量。研究还建立七维分类体系,识别三类解释不公根源(表示驱动、解释模型错配、可行动性驱动),并设计六步可操作审计流程,推动解释公平性成为负责任AI的核心指标。
原文摘要 · Abstract (English)
Machine learning algorithms are being used in high-stakes decisions, including those in criminal justice, healthcare, credit, and employment. The research community has responded with two largely independent research fields: \emph{algorithmic fairness}, which targets equitable outcomes, and \emph{explainable AI} (XAI), which targets interpretable reasoning. This survey identifies and maps a novel blind spot at their intersection, which is a model that can satisfy every standard fairness criterion in its outputs while being profoundly unfair in its \emph{reasoning process}. We refer to this as the procedural bias, and mitigating it requires treating the fairness of explanations as a distinct object of scientific study. To our knowledge, we provide the first unified theoretical and literature review of this emerging field and elucidate the drawbacks of post-hoc explainers in certifying explanation fairness. Our central contribution is a \emph{conditional invariance framework} formalizing explanation fairness as the requirement that explanations should be indifferent regardless of the protected attributes $ P(E(X) \in \cdot \mid X_\text{rel} = x_\text{rel},\, A = a) = P(E(X) \in \cdot \mid X_\text{rel} = x_\text{rel},\, A = b)$ for all task-relevant $x$, a single principle from which all existing explanation fairness metrics emerge as partial operationalizations. We introduce a seven-dimensional taxonomy, identify three generative mechanisms of explanation inequity (representation-driven, explanation-model mismatch, actionability-driven), and propose a canonical six-step evaluation workflow for operationalizing explanation fairness audits in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。