arXiv:2605.28464cs.CLcs.AI2026-05

提出新任务PDP,让AI预测案件是否起诉,补全法律责任评估盲区。

The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment

论文配图:The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment
图 1 · 摘自论文原文
  • 构建基于检察机关审查的案件分类任务PDP,涵盖起诉与三类不诉决定
  • 在4630个真实中国案件上测试,主流大模型表现远逊于传统LJP任务
  • 揭示现有AI在证据判断与价值权衡上能力不足,适合法律AI研究者参考

法律判决预测(LJP)已成为评估AI在刑事法律领域表现的核心基准,但仅覆盖已通过检察审查并正式起诉的案件,导致对证据不足、无刑事责任或免于处罚的案件存在重大盲区。为填补这一空白,我们提出首个以检察审查为核心的法律AI任务——起诉决策预测(PDP),将案件分类为起诉或三类不诉决定,反映AI在证据评估、法律归入和价值判断方面的能力。我们进一步构建了包含4,630个真实中国检察机关决定的PDP-Bench基准,涵盖190项罪名。大量实验表明,当前先进大模型在PDP任务上的表现显著低于其在LJP任务中的表现,且主流增强方法无法缩小差距。此外,受控的强化学习干预显示,简单的结果奖励无法生成可泛化的起诉决策能力。

原文摘要 · Abstract (English)

Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have already passed prosecutorial review and been formally indicted. As a result, LJP leaves a substantial blind spot in assessing criminal liability, overlooking cases involving insufficient evidence, no criminal liability, or guilt exempted from punishment. To fill this gap, we propose \textbf{Prosecution Decision Prediction (PDP)}, the first Legal AI task built around prosecutorial review, which classifies each case into prosecution or one of three non-prosecution decisions and reflects legal AI's capabilities in evidence evaluation, legal subsumption, and value-based discretion. We further construct \textbf{PDP-Bench}, a benchmark of 4{,}630 real Chinese prosecutorial decisions spanning 190 charges. Extensive experiments show that state-of-the-art LLMs perform substantially worse on PDP than on LJP and that mainstream enhancement routes fail to close the gap. Moreover, controlled RLVR interventions show that simple outcome rewards fail to produce generalizable PDP discrimination.

法律AI起诉预测司法数据大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。