研究监释再犯风险评估中模型多样性导致的预测随意性问题。
Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment
- 构建千名囚犯释放数据集,训练可解释模型提升预测性能。
- 实证发现相似准确模型间预测一致性高于理论下界。
- 采用最低风险评分策略可有效缓解预测随意性,适合高风险决策场景。
针对个体未来预测任务固有的噪声特性,当多个表现相近的模型对同一人给出不同预测时,会引发决策任意性问题。我们以一个已使用超过15年的再犯风险评估机器学习辅助系统为研究对象,将复杂法律规则转化为算法标签(再犯/未再犯),构建了包含数千名囚犯释放的数据集。在此基础上,训练出可解释模型,提升了预测性能,降低了群体间误差率差异,并确保康复进展能降低风险评分。进一步研究预测多重性:首先推导出有限模型集在数据集上预期预测一致性的紧下界;随后评估结构多样性(如不同模型系数)是否导致预测差异。实验表明,尽管存在大量表现相近的模型,但其预测一致性远高于最坏情况理论预测。提出仅取各模型中最低风险评分的简单策略,能有效应对预测任意性。
原文摘要 · Abstract (English)
Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models. When these models produce different predictions for the same individual, they raise concerns of arbitrariness in decision-making. How severe can this arbitrariness be, in theory and in practice? How can it be resolved to support high-stakes risk assessment? We address these questions through a study of a machine learning-based decision support system for recidivism risk assessment that has been in use for over 15 years. By translating complex legal rules into an algorithm for labeling post release outcomes (recidivist or non-recidivist), we first construct a dataset of thousands of inmate releases. Using this dataset, we learn interpretable models that improve predictive performance, reduce error-rate disparities between groups, and ensure that rehabilitative progress lowers risk scores. Next, we study predictive multiplicity, by first deriving a tight lower bound on the expected predictive agreement of any finite set of models over a dataset, and then by evaluating the extent to which structural diversity (e.g., different model coefficients) within this set translates to predictive multiplicity (i.e., different predictions for the same individual). Our experiments indicate that the existence of many similarly accurate models with comparable error-rate disparities does not necessarily translate into severe predictive multiplicity. Empirically, similarly performant models can exhibit substantially higher predictive agreement than worst-case theoretical guarantees suggest. We find that a simple policy that assigns each inmate the lowest risk among these models is effective for addressing predictive arbitrariness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。