arXiv:2412.14588cs.CL2024-12EMNLP被引 10

首个支持无罪判决的法律判决预测数据集,提升AI判案精准度

Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning

  • 基于三段式刑法推理框架构建新数据集
  • 现有模型在无罪案件上F1低于0.3
  • 新方法显著提升无罪判决预测能力

在法律实践中,法官遵循刑法三段式要件——构成要件、违法性与有责性,依次判断行为是否构成犯罪。尽管当前法律大模型在判决预测上表现出一定准确性,但因缺乏合适的基准数据集,普遍不具备三段式推理能力,导致所有输入均被自动判定为有罪,无法预测无罪结果,限制了其实际应用价值。为此,我们提出LJPIV,首个支持无罪判决的法律判决预测基准数据集。严格遵循三段式法理逻辑,通过大模型增强与人工校验,扩展三个主流法律数据集。实验表明:(1)现有法律大模型在LJPIV上表现不佳,最佳模型F1分数不足0.3;(2)引入三段式推理的新策略,在零样本提示与微调中显著提升域内及跨域判决预测准确率,尤其改善无罪判决预测效果。

原文摘要 · Abstract (English)

In legal practice, judges apply the trichotomous dogmatics of criminal law, sequentially assessing the elements of the offense, unlawfulness, and culpability to determine whether an individual's conduct constitutes a crime. Although current legal large language models (LLMs) show promising accuracy in judgment prediction, they lack trichotomous reasoning capabilities due to the absence of an appropriate benchmark dataset, preventing them from predicting innocent outcomes. As a result, every input is automatically assigned a charge, limiting their practical utility in legal contexts. To bridge this gap, we introduce LJPIV, the first benchmark dataset for Legal Judgment Prediction with Innocent Verdicts. Adhering to the trichotomous dogmatics, we extend three widely-used legal datasets through LLM-based augmentation and manual verification. Our experiments with state-of-the-art legal LLMs and novel strategies that integrate trichotomous reasoning into zero-shot prompting and fine-tuning reveal: (1) current legal LLMs have significant room for improvement, with even the best models achieving an F1 score of less than 0.3 on LJPIV; and (2) our strategies notably enhance both in-domain and cross-domain judgment prediction accuracy, especially for cases resulting in an innocent verdict.

法律AI三段式推理无罪判决数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。