处理保释决策中无法观测的反事实结果,避免模型引入偏见。
Confronting Label Indeterminacy in Automated Bail Decisions

- 提出新方法模拟未获保释者是否出庭的潜在结果
- 五种处理方式均影响模型预测,部分影响超过模型选择本身
- 揭示算法决策过程受标签不确定性显著干扰,适合司法AI研究者
保释决策面临数据驱动系统的核心挑战:当保释被拒时,被告是否出庭的反事实结果无法观测,导致历史数据存在结构性标签不确定性。未来决策受过去决策影响,而这些决策的结果仅部分可知,基于此类数据构建自动化系统可能引入偏见并形成反馈循环。本文以宾夕法尼亚州统一司法系统数据为案例,评估五种应对标签不确定性的现代方法,涵盖三种机器学习模型及一种基于保释动态的新标签插补方法。每种方法均依赖不可验证假设,但均显著影响模型预测行为,甚至超越模型选择的影响。可解释AI分析进一步显示,这些影响延伸至模型内部决策机制。最后从法律视角审视标签不确定性,评估各方法在保释决策中的正当性。
原文摘要 · Abstract (English)
Bail decisions present a fundamental challenge for data-driven decision support systems. When bail is denied, the counterfactual outcome of whether the defendant would have appeared in court remains unobserved. As a result, historical bail data embed structural label indeterminacy: future decisions are influenced by past decisions whose outcomes are only partially knowable. Building automated systems on such data risks introducing bias and reinforcing feedback loops. This raises a core question for machine-learning systems intended to assist judicial actors: how should cases in which bail was denied be treated during model development? In a case study of bail decisions from the Unified Judicial System of Pennsylvania, we evaluate five contemporary approaches to handling label indeterminacy across three machine learning models, including a novel label imputation method motivated by the dynamics of bail decisions. Each method relies on unverifiable assumptions, yet all influence the models' predictive behaviour, sometimes even more so than the choice of model itself. Explainable AI analysis further reveals that these effects extend to the models' internal decision-making processes as well. Finally, we consider the notion of label indeterminacy from a legal perspective and assess the legitimacy of these approaches in the context of bail decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。