用部分识别法评估算法公平性,从数据中推断反事实公平的可信范围。
Partial Identification Approach to Counterfactual Fairness Assessment
- 基于观测数据构建反事实公平性的置信区间,不依赖模型内部结构
- 在COMPAS数据上发现种族变更导致评分虚高,年龄增长则真实降低评分
- 适合关注算法公平性但缺乏完整数据的研究者与政策制定者
AI决策系统在刑事司法、贷款审批和招聘等关键领域广泛应用,引发对算法公平性的关注。由于通常只能获取算法输出而无法了解其内部机制,研究者常通过改变辅助敏感属性(如种族)来考察决策变化,进而提出反事实公平性度量。然而,如何从现有数据评估该度量仍具挑战性。许多实际场景中,目标反事实度量不可识别,即无法由定量数据与定性知识唯一确定。本文采用部分识别方法,从观测数据中推导出反事实公平度量的有信息量边界。我们引入贝叶斯方法,以高置信度界定未知的反事实公平度量。我们在COMPAS数据集上验证算法,评估了种族、年龄和性别对再犯风险评分的公平性影响。结果表明,将种族变更为非裔美国人时,评分呈现正向(虚假)效应;从年轻变为年老时,则呈现负向(直接因果)效应。
原文摘要 · Abstract (English)
The wide adoption of AI decision-making systems in critical domains such as criminal justice, loan approval, and hiring processes has heightened concerns about algorithmic fairness. As we often only have access to the output of algorithms without insights into their internal mechanisms, it was natural to examine how decisions would alter when auxiliary sensitive attributes (such as race) change. This led the research community to come up with counterfactual fairness measures, but how to evaluate the measure from available data remains a challenging task. In many practical applications, the target counterfactual measure is not identifiable, i.e., it cannot be uniquely determined from the combination of quantitative data and qualitative knowledge. This paper addresses this challenge using partial identification, which derives informative bounds over counterfactual fairness measures from observational data. We introduce a Bayesian approach to bound unknown counterfactual fairness measures with high confidence. We demonstrate our algorithm on the COMPAS dataset, examining fairness in recidivism risk scores with respect to race, age, and sex. Our results reveal a positive (spurious) effect on the COMPAS score when changing race to African-American (from all others) and a negative (direct causal) effect when transitioning from young to old age.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。