arXiv:2603.13452cs.AIcs.CY2026-03

提出新指标检测模型在交叉群体中的解释差异,发现传统公平评估遗漏的问题。

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

  • 基于多属性交叉构建解释稳定性差距度量,融合条件风险加权与贝叶斯收缩
  • 在三个数据集上验证,发现4种主流方法存在被忽视的交叉群体解释不公平
  • 适合关注模型可解释性、监管合规及交叉歧视问题的研究者使用

机器学习公平性评估主要依赖结果导向指标(如人口统计均等),但无法检测模型对不同群体是否采用系统性不同的推理方式,违背程序公平原则。这一问题在交叉性(intersectionality)下尤为严重:模型在单一属性(如种族)上看似公平,却在交叉子群体(如种族×性别)中出现显著差异,即公平性操纵现象。本文提出多类别解释稳定性差异(MESD),量化由多个受保护属性的笛卡尔积形成的交叉子群体间解释质量的差异。MESD整合三要素:与结果条件一致的标签感知聚合、用于小群体估计稳定的经验贝叶斯收缩,以及强调最差情况子群体差异的条件风险价值(CVaR)加权。我们将MESD嵌入多目标优化框架UEF,联合优化效用、结果公平性和程序公平性,使用NSGA-II算法。在三个基准数据集和四种前沿方法上进行多组实验,结果表明MESD能揭示仅靠结果指标无法察觉的程序性不公平。研究将贡献置于程序正义理论框架内,并讨论其对法规合规与交叉平等的意义。

原文摘要 · Abstract (English)

Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predictions are statistically consistent across protected groups. However, these metrics cannot detect whether a model uses systematically different reasoning for different demographic groups, which violates procedural fairness principles. This problem is compounded by intersectionality, where models may appear fair on individual attributes (e.g., race) while exhibiting significant disparities for intersectional subgroups (e.g., race $\times$ gender), a phenomenon known as fairness gerrymandering. In this work, we introduce Multi-category Explanation Stability Disparity (MESD), a procedural fairness metric that quantifies disparities in explanation quality across intersectional subgroups formed by the Cartesian product of multiple protected attributes. MESD integrates three components, which are label-aware aggregation aligned with outcome-conditional fairness, empirical-Bayes shrinkage to stabilize estimates for small intersectional groups, and Conditional Value-at-Risk (CVaR) weighting to emphasize worst-case subgroup disparities. We integrate MESD within a multi-objective optimization framework (UEF) that jointly optimizes utility, outcome fairness, and procedural fairness using NSGA-II. We evaluated MESD and UEF on three benchmark datasets along with four state-of-the-art methods in several experiments, and we demonstrate that MESD reveals procedural disparities invisible to outcome metrics alone. We position our contribution within procedural justice theory and discuss implications for regulatory compliance and intersectional equity.

公平性评估交叉性解释公平程序公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。