arXiv:2502.01267cs.LG2025-02被引 5

通过反事实推理检测模型中个体歧视,支持单维度与多维度分析。

Counterfactual Situation Testing: From Single to Multidimensional Discrimination

  • 基于反事实生成被投诉者若属性改变后的虚拟样本,替代传统相似样本对比。
  • 实验显示其比传统方法发现更多歧视案例,即使模型已满足反事实公平。
  • 适用于法律合规审查、算法审计,尤其适合关注交叉歧视的研究者。

本文提出反事实情境测试(CST),一种用于检测分类器决策数据集中个体歧视的因果数据挖掘框架。CST 回答的问题是:如果当事人具有不同的受保护属性,模型结果会如何?它扩展了法律基础的情境测试(ST),通过反事实推理实现“给定差异下的公平性”。传统 ST 为每位投诉人寻找同类受保护与非受保护实例,构建对照组与测试组进行比较;而 CST 则避免理想化匹配,转而以被投诉者的反事实生成样本作为测试组,反映当受保护属性变化时其他看似中立属性的影响。在单维(如性别)与多维(如性别与种族)歧视检测中,我们发现多维歧视下交集歧视常被忽略。采用 k-近邻实现,我们在合成与真实数据上验证了 CST。实验表明,即使模型满足反事实公平,CST 仍能识别出更多歧视案例。CST 进一步扩展了 Kusner 等人提出的反事实公平(CF),为其引入置信区间,并在所有实验中报告该区间。

原文摘要 · Abstract (English)

We present counterfactual situation testing (CST), a causal data mining framework for detecting individual discrimination in a dataset of classifier decisions. CST answers the question ``what would have been the model outcome had the individual, or complainant, been of a different protected status?'' It extends the legally-grounded situation testing (ST) of Thanh et al. (2011) by operationalizing the notion of "fairness given the difference" via counterfactual reasoning. ST finds for each complainant similar protected and non-protected instances in the dataset; constructs, respectively, a control and test group; and compares the groups such that a difference in model outcomes implies a potential case of individual discrimination. CST, instead, avoids this idealized comparison by establishing the test group on the complainant's generated counterfactual, which reflects how the protected attribute when changed influences other seemingly neutral attributes of the complainant. Under CST we test for discrimination for each complainant by comparing similar individuals within the control and test group but dissimilar individuals across these groups. We consider single (e.g.,~gender) and multidimensional (e.g.,~gender and race) discrimination testing. For multidimensional discrimination we study multiple and intersectional discrimination and, as feared by legal scholars, find evidence that the former fails to account for the latter kind. Using a k-nearest neighbor implementation, we showcase CST on synthetic and real data. Experimental results show that CST uncovers a higher number of cases than ST, even when the model is counterfactually fair. CST, in fact, extends counterfactual fairness (CF) of Kusner et al. (2017) by equipping CF with confidence intervals, which we report for all experiments.

反事实推理歧视检测公平性多维歧视

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。