arXiv:2502.01211cs.LGstat.ML2025-02被引 1

用特权得分量化偏见,帮机器学习更公平。

Privilege Scores

  • 通过真实与公平世界对比,量化保护属性带来的特权
  • 能识别需平权干预的个体,也可指导政策制定
  • 提供可解释性工具,揭示特权来源的中介特征

公平感知机器学习中的偏见转化方法旨在纠正受保护属性(PA)影响下的非中立现状。然而现有方法缺乏对非中立成因的明确表述。本文提出特权得分(PS),通过比较真实世界与移除PA影响后的公平世界中的模型预测,衡量PA相关的特权。在个体层面,PS可识别适合平权行动的对象;在全局层面,可用于指导偏见转化政策。我们提出了PS的估计方法,并引入特权得分贡献(PSCs)作为解释工具,归因于中介特征和直接效应。同时提供PS与PSCs的置信区间。在模拟和真实数据上的实验验证了方法的广泛应用性,并对房贷与大学录取中的性别和种族特权提供了新见解。

原文摘要 · Abstract (English)

Bias-transforming methods of fairness-aware machine learning aim to correct a non-neutral status quo with respect to a protected attribute (PA). Current methods, however, lack an explicit formulation of what drives non-neutrality. We introduce privilege scores (PS) to measure PA-related privilege by comparing the model predictions in the real world with those in a fair world in which the influence of the PA is removed. At the individual level, PS can identify individuals who qualify for affirmative action; at the global level, PS can inform bias-transforming policies. After presenting estimation methods for PS, we propose privilege score contributions (PSCs), an interpretation method that attributes the origin of privilege to mediating features and direct effects. We provide confidence intervals for both PS and PSCs. Experiments on simulated and real-world data demonstrate the broad applicability of our methods and provide novel insights into gender and racial privilege in mortgage and college admissions applications.

公平性偏见量化可解释性平权算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。