提出统计检验框架,识别量刑预测算法中真实存在的不公平偏差。
Detecting Statistically Significant Fairness Violations in Recidivism Forecasting Algorithms
- 用k折交叉验证生成公平性指标的采样分布,进行统计显著性检验。
- 在司法数据上发现黑人群体存在显著不公平,白人则无或反向偏差。
- 适合关注算法公平性评估的政策制定者与研究人员使用。
机器学习算法正被广泛应用于金融、医疗和刑事司法等关键领域。学术界对算法公平性日益关注,已有研究提出多种公平性定义,用于量化特权群体与受保护群体间的差异,并通过因果推断分析种族对模型预测的影响,以及检验模型概率预测的校准性。然而,现有文献缺乏判断观察到的群体差异是否具有统计显著性,还是仅由随机因素导致的方法。本文提出一种严格的框架,利用k折交叉验证生成公平性指标的采样分布,进而实施统计检验,以识别基于预测与实际结果差异、模型校准性及因果推断技术的公平性违规是否具有统计显著性。我们在国家司法研究所的数据上测试了量刑再犯预测算法,结果表明:在多个公平性定义下,针对黑人群体存在显著偏差;而在其他定义下,则未发现偏差或呈现对白人群体的偏差。研究强调了在评估算法决策系统时,必须采用严谨且稳健的统计检验方法。
原文摘要 · Abstract (English)
Machine learning algorithms are increasingly deployed in critical domains such as finance, healthcare, and criminal justice [1]. The increasing popularity of algorithmic decision-making has stimulated interest in algorithmic fairness within the academic community. Researchers have introduced various fairness definitions that quantify disparities between privileged and protected groups, use causal inference to determine the impact of race on model predictions, and that test calibration of probability predictions from the model. Existing literature does not provide a way in which to assess whether observed disparities between groups are statistically significant or merely due to chance. This paper introduces a rigorous framework for testing the statistical significance of fairness violations by leveraging k-fold cross-validation [2] to generate sampling distributions of fairness metrics. This paper introduces statistical tests that can be used to identify statistically significant violations of fairness metrics based on disparities between predicted and actual outcomes, model calibration, and causal inference techniques [1]. We demonstrate this approach by testing recidivism forecasting algorithms trained on data from the National Institute of Justice. Our findings reveal that machine learning algorithms used for recidivism forecasting exhibit statistically significant bias against Black individuals under several fairness definitions, while also exhibiting no bias or bias against White individuals under other definitions. The results from this paper underscore the importance of rigorous and robust statistical testing while evaluating algorithmic decision-making systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。