LegalOne通过三阶段训练,让大模型学会精准可靠的中文法律推理。
LegalOne: A Family of Foundation Models for Reliable Legal Reasoning
- 用动态采样策略平衡学新知识和保留原能力,打牢法律基础
- 通过代理式思维链提炼,使模型推理过程结构化、事实准确
- 分阶段强化学习,实现从记忆到自主推理的跃迁,适合司法场景
尽管大型语言模型展现出强大的通用能力,但其在法律领域的直接应用常因缺乏精确领域知识以及复杂多步司法推理能力而受限。为此,我们提出LegalOne,一个专为中文法律领域设计的基础模型家族。LegalOne通过三阶段全流程训练实现法律推理能力的掌握:第一阶段,采用基于困惑度的可塑性调整采样(PAS),在获取新知识与保留原始能力间取得平衡,建立坚实的法律基础;第二阶段,在监督微调中使用法律代理思维链蒸馏(LEAD),将原始法律文本中的复杂司法流程转化为结构化的推理路径,确保事实依据和逻辑严谨性;第三阶段,实施课程强化学习(RL),通过从记忆、理解到推理的渐进式强化过程,使模型从简单模式匹配进化为自主且可靠的法律推理。实验表明,LegalOne在多项法律任务上达到顶尖水平,超越参数量大得多的通用大模型,凭借更高的知识密度与效率。我们公开发布LegalOne权重与LegalKit评估框架,推动法律AI发展,为高风险司法应用部署可信、可解释的基础模型铺平道路。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated impressive general capabilities, their direct application in the legal domain is often hindered by a lack of precise domain knowledge and complexity of performing rigorous multi-step judicial reasoning. To address this gap, we present LegalOne, a family of foundational models specifically tailored for the Chinese legal domain. LegalOne is developed through a comprehensive three-phase pipeline designed to master legal reasoning. First, during mid-training phase, we propose Plasticity-Adjusted Sampling (PAS) to address the challenge of domain adaptation. This perplexity-based scheduler strikes a balance between the acquisition of new knowledge and the retention of original capabilities, effectively establishing a robust legal foundation. Second, during supervised fine-tuning, we employ Legal Agentic CoT Distillation (LEAD) to distill explicit reasoning from raw legal texts. Unlike naive distillation, LEAD utilizes an agentic workflow to convert complex judicial processes into structured reasoning trajectories, thereby enforcing factual grounding and logical rigor. Finally, we implement a Curriculum Reinforcement Learning (RL) strategy. Through a progressive reinforcement process spanning memorization, understanding, and reasoning, LegalOne evolves from simple pattern matching to autonomous and reliable legal reasoning. Experimental results demonstrate that LegalOne achieves state-of-the-art performance across a wide range of legal tasks, surpassing general-purpose LLMs with vastly larger parameter counts through enhanced knowledge density and efficiency. We publicly release the LegalOne weights and the LegalKit evaluation framework to advance the field of Legal AI, paving the way for deploying trustworthy and interpretable foundation models in high-stakes judicial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。