arXiv:2604.27282cs.CYcs.LG2026-04中稿 · the 2026 ACM Confe…

低发暴力再犯率下,现有评估工具难以准确识别高风险者。

The Likelihood Ratio Wall: Structural Limits on Accurate Risk Assessment for Rare Violence

  • 提出似然比墙理论,揭示低再犯率下工具区分能力的固有极限
  • 当再犯率2%-5%时,高风险标签者中真正再犯者不足一半
  • 提醒决策者:应公开说明评估结果的不确定性,避免误用

每年有超过一百万美国被告使用保释风险评估工具,但其预测罕见暴力再犯面临基本统计障碍。我们推导出一个通用精度上限——似然比墙,表明当暴力再逮捕率仅为2%-5%时,即使达到50%的命中率(阳性预测值,PPV),也需远超当前工具的区分能力。对罕见事件而言,工具可能看似表现良好,却在大多数被标记为“高风险”时出错。事后评分校准无法解决此问题,因其不提升工具分离真阳与假阳的本质能力。我们进一步证明了监控天花板:过度执法使未再犯者出现更多记录“风险因素”,导致被过度监控群体的最大可实现精度更低,即使实际犯罪率相同。我们将结果转化为需羁押人数(Number Needed to Detain),以防止一次暴力犯罪需羁押多少人,并建议风险报告应明确传达这种不确定性。研究指出,在当前数据条件下,仅讨论公平性指标是不完整的;可用特征可能不足以支持高置信度的个体化监禁决策。

原文摘要 · Abstract (English)

Pretrial risk assessment tools are used on over one million U.S. defendants each year, yet their use for predicting rare violent re-offense faces a basic statistical barrier. We derive a universal precision bound -- the Likelihood Ratio Wall -- showing that when violent re-arrest rates are low (2-5%), achieving even a 50% hit rate among people labeled "high risk" (positive predictive value, or PPV) would require tools far more discriminative than current instruments appear to be. For rare outcomes, a tool can have respectable-looking performance metrics and still be wrong most of the time it flags someone as "high risk for violence." We show that post-hoc score recalibration cannot solve this problem because it does not improve the tool's underlying ability to separate true positives from false positives. We further prove a Surveillance Ceiling: when over-policing inflates recorded "risk factors" among those who would not re-offend, the maximum achievable precision is structurally lower for over-policed groups, even at equal offense rates. We translate these results into the Number Needed to Detain (how many people must be detained to prevent one violent offense), and propose that risk reports should communicate this uncertainty explicitly. Our findings suggest that for rare violent outcomes, debates about fairness metrics alone are incomplete: under current data regimes, the available features may not support high-confidence individualized detention decisions.

风险评估再犯预测统计局限司法公正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。