arXiv:2601.01331cs.CYcs.CL2026-01被引 1

构建上诉案判决生成基准,推动法律大模型理解二审推理逻辑

AppellateGen: A Benchmark for Appellate Legal Judgment Generation

  • 设计基于司法流程的多智能体系统,分阶段生成上诉判决
  • 包含7351对案例,需结合初审结果与新证据推理
  • 适合法律AI研究者及司法智能化开发者参考

法律判决生成是法律智能的关键任务。现有研究多聚焦于一审案件,依赖静态事实到判决的映射,忽视了上诉(二审)审查的辩证性。为此,我们提出AppellateGen,一个包含7351对案例的二审判决生成基准,要求模型基于初审判决和证据更新进行推理,以建模审判阶段间的因果关系。我们进一步提出基于司法标准操作流程的法律多智能体系统(SLMAS),将生成过程分解为问题识别、检索和起草等离散阶段。实验表明,尽管SLMAS提升了逻辑一致性,当前大模型在上诉推理复杂性面前仍面临重大挑战。数据集与代码已公开:https://anonymous.4open.science/r/AppellateGen-5763。

原文摘要 · Abstract (English)

Legal judgment generation is a critical task in legal intelligence. However, existing research in legal judgment generation has predominantly focused on first-instance trials, relying on static fact-to-verdict mappings while neglecting the dialectical nature of appellate (second-instance) review. To address this, we introduce AppellateGen, a benchmark for second-instance legal judgment generation comprising 7,351 case pairs. The task requires models to draft legally binding judgments by reasoning over the initial verdict and evidentiary updates, thereby modeling the causal dependency between trial stages. We further propose a judicial Standard Operating Procedure (SOP)-based Legal Multi-Agent System (SLMAS) to simulate judicial workflows, which decomposes the generation process into discrete stages of issue identification, retrieval, and drafting. Experimental results indicate that while SLMAS improves logical consistency, the complexity of appellate reasoning remains a substantial challenge for current LLMs. The dataset and code are publicly available at: https://anonymous.4open.science/r/AppellateGen-5763.

法律AI判决生成多智能体上诉裁判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。