构建民事上诉案件数据集,推动法律AI研究从判决分析转向上诉审查。
AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios
- 构建1万对真实一审与二审文书数据,覆盖91类民事案件。
- 现有模型在判断是否改判任务上F1低于50%,凸显挑战性。
- 适合法律AI、司法自动化研究者,助力提升裁判一致性。
近年来的LegalAI研究多集中于个案判决分析,忽视了司法体系中的关键上诉环节。上诉是纠错与保障公正审判的核心机制,具有重要实践与研究价值。为此,我们提出AppealCase数据集,包含10,000对真实世界中匹配的一审与二审文书,涵盖91类民事案件。数据集附有五个维度的详细标注:判决改判情况、改判原因、援引法律条文、诉求层面的裁决结果,以及二审中是否存在新信息。基于此,我们提出五项新的LegalAI任务,并在20个主流模型上进行综合评估。实验结果表明,所有模型在判决改判预测任务上的F1分数均低于50%,凸显上诉场景的复杂性与挑战性。我们希望AppealCase能推动LegalAI在上诉分析领域的研究,促进司法决策的一致性提升。
原文摘要 · Abstract (English)
Recent advances in LegalAI have primarily focused on individual case judgment analysis, often overlooking the critical appellate process within the judicial system. Appeals serve as a core mechanism for error correction and ensuring fair trials, making them highly significant both in practice and in research. To address this gap, we present the AppealCase dataset, consisting of 10,000 pairs of real-world, matched first-instance and second-instance documents across 91 categories of civil cases. The dataset also includes detailed annotations along five dimensions central to appellate review: judgment reversals, reversal reasons, cited legal provisions, claim-level decisions, and whether there is new information in the second instance. Based on these annotations, we propose five novel LegalAI tasks and conduct a comprehensive evaluation across 20 mainstream models. Experimental results reveal that all current models achieve less than 50% F1 scores on the judgment reversal prediction task, highlighting the complexity and challenge of the appeal scenario. We hope that the AppealCase dataset will spur further research in LegalAI for appellate case analysis and contribute to improving consistency in judicial decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。