arXiv:2510.10072cs.CL2025-10EMNLP被引 14

70亿参数模型用强化学习提升法律推理能力,效果媲美更大模型。

Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference

  • 基于17000条法律推理链数据,分两阶段训练增强逻辑与知识。
  • 在权威评测中表现超越同规模模型,接近540亿参数大模型水平。
  • 适合法律AI研发者、司法科技从业者快速部署高可靠推理系统。

当前聚焦推理的大语言模型在多个领域快速发展,但在复杂法律问题上的能力仍待探索。本文提出面向法律推理的轻量级模型Unilaw-R1,仅含70亿参数,显著降低部署成本,有效应对法律知识不足、推理逻辑不可靠和业务泛化能力弱三大挑战。我们构建了包含1.7万条提炼筛选后的思维链(CoT)样本的高质量数据集Unilaw-R1-Data。在此基础上,采用监督微调(SFT)与强化学习(RL)相结合的两阶段训练策略,大幅提升了复杂法律推理任务表现,并支持可解释的决策过程。为评估法律推理能力,我们设计了专用基准Unilaw-R1-Eval,涵盖单选与多选任务。实验表明,Unilaw-R1在主流基准上表现优异,超越所有同规模模型,性能接近540亿参数的DeepSeek-R1-Distill-Qwen-32B(54.9%)。经领域特定训练后,在LawBench和LexEval上平均超过Qwen-2.5-7B-Instruct(46.6%)6.6个百分点。

原文摘要 · Abstract (English)

Reasoning-focused large language models (LLMs) are rapidly evolving across various domains, yet their capabilities in handling complex legal problems remains underexplored. In this paper, we introduce Unilaw-R1, a large language model tailored for legal reasoning. With a lightweight 7-billion parameter scale, Unilaw-R1 significantly reduces deployment cost while effectively tackling three core challenges in the legal domain: insufficient legal knowledge, unreliable reasoning logic, and weak business generalization. To address these issues, we first construct Unilaw-R1-Data, a high-quality dataset containing 17K distilled and screened chain-of-thought (CoT) samples. Based on this, we adopt a two-stage training strategy combining Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), which significantly boosts the performance on complex legal reasoning tasks and supports interpretable decision-making in legal AI applications. To assess legal reasoning ability, we also introduce Unilaw-R1-Eval, a dedicated benchmark designed to evaluate models across single- and multi-choice legal tasks. Unilaw-R1 demonstrates strong results on authoritative benchmarks, outperforming all models of similar scale and achieving performance on par with the much larger DeepSeek-R1-Distill-Qwen-32B (54.9%). Following domain-specific training, it also showed significant gains on LawBench and LexEval, exceeding Qwen-2.5-7B-Instruct (46.6%) by an average margin of 6.6%.

法律AI推理模型强化学习小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。