arXiv:2502.17943cs.CL2025-02被引 12

首个中文法律文书多阶段生成基准,助力AI写判决书。

CaseGen: A Benchmark for Multi-Stage Legal Case Documents Generation

  • 基于500个真实案件构建多阶段法律文书生成任务
  • 涵盖7个核心部分,支持4类关键写作任务
  • 首创LLM作为裁判的评估框架,适合法律AI研究者

法律文书在司法程序中至关重要。随着案件数量持续上升,人工撰写法律文书面临巨大压力。大语言模型(LLMs)为自动化文档生成提供了可能,但现有基准未能充分反映真实场景中的复杂性。为此,我们提出CaseGen,首个面向中文法律领域多阶段文书生成的基准。该数据集基于500个经法律专家标注的真实案件样本,覆盖7个核心文书部分,支持4项关键任务:辩护意见起草、庭审事实撰写、法律推理构建和判决结果生成。据我们所知,CaseGen是首个专为评估LLMs在法律文书生成中表现而设计的基准。为确保评估准确性,我们设计了基于LLM作为裁判的评估框架,并通过人工标注验证其有效性。我们评估了多个通用领域与法律专用的LLMs,揭示其在文书生成中的局限性并指明改进方向。本工作推动了法律文书自动化的有效框架建设,为AI在法律领域的可靠应用铺路。数据与代码已公开于https://github.com/CSHaitao/CaseGen。

原文摘要 · Abstract (English)

Legal case documents play a critical role in judicial proceedings. As the number of cases continues to rise, the reliance on manual drafting of legal case documents is facing increasing pressure and challenges. The development of large language models (LLMs) offers a promising solution for automating document generation. However, existing benchmarks fail to fully capture the complexities involved in drafting legal case documents in real-world scenarios. To address this gap, we introduce CaseGen, the benchmark for multi-stage legal case documents generation in the Chinese legal domain. CaseGen is based on 500 real case samples annotated by legal experts and covers seven essential case sections. It supports four key tasks: drafting defense statements, writing trial facts, composing legal reasoning, and generating judgment results. To the best of our knowledge, CaseGen is the first benchmark designed to evaluate LLMs in the context of legal case document generation. To ensure an accurate and comprehensive evaluation, we design the LLM-as-a-judge evaluation framework and validate its effectiveness through human annotations. We evaluate several widely used general-domain LLMs and legal-specific LLMs, highlighting their limitations in case document generation and pinpointing areas for potential improvement. This work marks a step toward a more effective framework for automating legal case documents drafting, paving the way for the reliable application of AI in the legal field. The dataset and code are publicly available at https://github.com/CSHaitao/CaseGen.

法律AI文本生成基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。