LexPro-1.0是专为中文法律场景打造的推理模型,提升法律任务精准度。
LexPro-1.0 Technical Report
- 基于20类罪名、31省份的数百万法律文书训练,强化领域适配性。
- 采用监督微调与无监督强化学习,显著提升推理能力与可解释性。
- 适合作为司法辅助工具,尤其适合需要严谨逻辑的复杂案件处理。
本文介绍首个面向中文法律领域的推理模型LexPro-1.0,旨在满足多样化的实际法律需求。现有法律大模型主要存在两大问题:其一,设计与评估多从计算机科学视角出发,缺乏法律专业知识与逻辑支持,难以胜任高精度法律任务;其二,因法律领域数据不足,导致模型在真实场景中表现受限。为此,我们收集了涵盖20余类犯罪、来自中国31个省份的数百万法律文档用于训练,并从中筛选高质量样本进行监督微调,确保内容相关性与准确性。模型进一步通过大规模无监督强化学习优化,重点提升推理能力与可解释性。为验证其在复杂法律应用中的有效性,我们邀请法律专家开展人工评估。基于DeepSeek-R1-Distilled模型,推出三种稠密配置版本:14B、32B和70B。
原文摘要 · Abstract (English)
In this report, we introduce our first-generation reasoning model, LexPro-1.0, a large language model designed for the highly specialized Chinese legal domain, offering comprehensive capabilities to meet diverse realistic needs. Existing legal LLMs face two primary challenges. Firstly, their design and evaluation are predominantly driven by computer science perspectives, leading to insufficient incorporation of legal expertise and logic, which is crucial for high-precision legal applications, such as handling complex prosecutorial tasks. Secondly, these models often underperform due to a lack of comprehensive training data from the legal domain, limiting their ability to effectively address real-world legal scenarios. To address this, we first compile millions of legal documents covering over 20 types of crimes from 31 provinces in China for model training. From the extensive dataset, we further select high-quality for supervised fine-tuning, ensuring enhanced relevance and precision. The model further undergoes large-scale reinforcement learning without additional supervision, emphasizing the enhancement of its reasoning capabilities and explainability. To validate its effectiveness in complex legal applications, we also conduct human evaluations with legal experts. We develop fine-tuned models based on DeepSeek-R1-Distilled versions, available in three dense configurations: 14B, 32B, and 70B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。