首个中文多步法律推理数据集,验证大模型法律推断能力
Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
- 基于真实判决书构建结构化法律推理数据集
- 模型自主生成思维链显著提升推理质量
- 适合法律AI与大模型推理研究者参考
大型语言模型在多个专业领域展现出强大的推理能力,推动其在法律推理中的应用研究。然而,现有法律评测基准常混淆事实记忆与真实推理,割裂推理过程,忽视推理质量。为此,我们提出MSLR,首个基于真实司法决策的中文多步法律推理数据集。MSLR采用IRAC框架(问题、规则、适用、结论)建模官方法律文件中的结构化专家推理。同时,设计可扩展的人机协同标注流程,高效生成细粒度步骤级推理标注,提供多步推理数据集的通用方法论。对多种大模型在MSLR上的评估显示性能仅达中等,凸显其在复杂法律推理任务中的适应挑战。进一步实验表明,由模型自主生成的自启动思维链提示优于人工设计提示,显著提升推理连贯性与质量。MSLR推动大模型推理与思维链策略发展,为后续研究提供开放资源。数据集与代码见https://github.com/yuwenhan07/MSLR-Bench和https://law.sjtu.edu.cn/flszyjzx/index.html。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong reasoning abilities across specialized domains, motivating research into their application to legal reasoning. However, existing legal benchmarks often conflate factual recall with genuine inference, fragment the reasoning process, and overlook the quality of reasoning. To address these limitations, we introduce MSLR, the first Chinese multi-step legal reasoning dataset grounded in real-world judicial decision making. MSLR adopts the IRAC framework (Issue, Rule, Application, Conclusion) to model structured expert reasoning from official legal documents. In addition, we design a scalable Human-LLM collaborative annotation pipeline that efficiently produces fine-grained step-level reasoning annotations and provides a reusable methodological framework for multi-step reasoning datasets. Evaluation of multiple LLMs on MSLR shows only moderate performance, highlighting the challenges of adapting to complex legal reasoning. Further experiments demonstrate that Self-Initiated Chain-of-Thought prompts generated by models autonomously improve reasoning coherence and quality, outperforming human-designed prompts. MSLR contributes to advancing LLM reasoning and Chain-of-Thought strategies and offers open resources for future research. The dataset and code are available at https://github.com/yuwenhan07/MSLR-Bench and https://law.sjtu.edu.cn/flszyjzx/index.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。