构建中文多被告多罪名判决预测数据集,揭示复杂案件对模型的挑战
MultiJustice: A Chinese Dataset for Multi-Party, Multi-Charge Legal Prediction
- 提出多被告多罪名法律判决预测新数据集MPMCP
- S4场景下模型性能下降最显著,F1最高损失达19.7%
- 适合法律AI研究者与司法大模型开发者参考
法律判决预测为法律从业者和研究人员提供了有力支持。然而,多个被告与多项罪名是否应分别处理仍缺乏研究。为此,我们构建了多被告多罪名预测数据集(MPMCP),并评估多种主流法律大模型在四种实际判决场景下的表现:(S1) 单被告单罪名,(S2) 单被告多罪名,(S3) 多被告单罪名,(S4) 多被告多罪名。在罪名预测与刑期预测两个任务上开展实验发现,S4场景最具挑战性,其次为S2、S3、S1。模型表现差异显著:相较于S1,InternLM2在S4中F1下降约4.5%,LogD上升2.8%;Lawformer则F1下降约19.7%,LogD上升19.0%。相关数据与代码已开源。
原文摘要 · Abstract (English)
Legal judgment prediction offers a compelling method to aid legal practitioners and researchers. However, the research question remains relatively under-explored: Should multiple defendants and charges be treated separately in LJP? To address this, we introduce a new dataset namely multi-person multi-charge prediction (MPMCP), and seek the answer by evaluating the performance of several prevailing legal large language models (LLMs) on four practical legal judgment scenarios: (S1) single defendant with a single charge, (S2) single defendant with multiple charges, (S3) multiple defendants with a single charge, and (S4) multiple defendants with multiple charges. We evaluate the dataset across two LJP tasks, i.e., charge prediction and penalty term prediction. We have conducted extensive experiments and found that the scenario involving multiple defendants and multiple charges (S4) poses the greatest challenges, followed by S2, S3, and S1. The impact varies significantly depending on the model. For example, in S4 compared to S1, InternLM2 achieves approximately 4.5% lower F1-score and 2.8% higher LogD, while Lawformer demonstrates around 19.7% lower F1-score and 19.0% higher LogD. Our dataset and code are available at https://github.com/lololo-xiao/MultiJustice-MPMCP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。