arXiv:2503.16040cs.CL2025-03EMNLP被引 14

对比多模型在法律推理中的表现,发现专精模型效果更优。

Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond

  • 用思维链扩展推理过程提升法律问答能力
  • 专有模型Legal-R1在多任务中表现接近顶尖水平
  • 适合法律AI研究者与司法科技开发者参考

近期大语言模型测试时扩展推理(test-time scaling)进展显著,如DeepSeek-R1和OpenAI o1通过延长思维链大幅提高通用推理性能。但其在法律推理领域的应用仍缺乏系统研究。为此,本文首次对12个模型(含推理专用与通用模型)在17项中英文法律任务上进行评估,覆盖成文法与判例法体系。我们通过蒸馏DeepSeek-R1构建双语法律思维链数据集,并推出开源模型Legal-R1,专攻法律领域。实验表明,Legal-R1在多任务中表现优异;DeepSeek-R1在中文法律任务中优势明显,o1在英文任务上表现相当。详细错误分析揭示常见问题:法律知识过时、解释力不足、事实幻觉频发。这些发现明确了法律类大模型的主要瓶颈,并指明未来研究方向。

原文摘要 · Abstract (English)

Recent advances in test-time scaling of large language models (LLMs), exemplified by DeepSeek-R1 and OpenAI's o1, show that extending the chain of thought during inference can significantly improve general reasoning performance. However, the impact of this paradigm on legal reasoning remains insufficiently explored. To address this gap, we present the first systematic evaluation of 12 LLMs, including both reasoning-focused and general-purpose models, across 17 Chinese and English legal tasks spanning statutory and case-law traditions. In addition, we curate a bilingual chain-of-thought dataset for legal reasoning through distillation from DeepSeek-R1 and develop Legal-R1, an open-source model specialized for the legal domain. Experimental results show that Legal-R1 delivers competitive performance across diverse tasks. DeepSeek-R1 exhibits clear advantages in Chinese legal reasoning, while OpenAI's o1 achieves comparable results on English tasks. We further conduct a detailed error analysis, which reveals recurring issues such as outdated legal knowledge, limited capacity for legal interpretation, and susceptibility to factual hallucinations. These findings delineate the main obstacles confronting legal-domain LLMs and suggest promising directions for future research.

法律AI推理增强开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。