专为政治法律领域打造的高效大模型,解决法律误引与推理弱问题。
PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs

- 融合持续预训练与强化学习,提升法律知识与推理能力。
- 在三个基准上表现优异,真实场景下效果最佳。
- 适合法律从业者、政策研究者及合规AI应用开发人员。
大语言模型在通用任务中表现卓越,但在法律领域仍面临法律引用错误、知识覆盖不全和结构化推理能力弱等挑战。为此,我们提出PoliLegalLM,一个针对政治与法律应用的领域专用大模型。通过统一训练框架,结合持续预训练、渐进式监督微调与基于偏好的强化学习,协同提升法律知识的准确性、任务对齐性与推理能力。我们构建了大规模高质量法律语料库,并设计结构化后训练流程,使模型有效学习领域知识并适应多样法律任务。在LawBench、LexEval及真实数据集PoliLegal上的评估显示,PoliLegalLM性能强劲且稳定,优于同规模竞争模型,甚至超越更大模型,在真实法律场景中表现最优。结果表明该训练范式有效,凸显领域专用大模型在实际法律应用中的价值。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable success in general-domain tasks, yet their direct application to the legal domain remains challenging due to hallucinated legal citations, incomplete knowledge coverage, and weak structured reasoning. To address these issues, we propose PoliLegalLM, a domain-specific large language model tailored for political and legal applications. Our approach adopts a unified training framework that integrates continued pretraining, progressive supervised fine-tuning, and preference-based reinforcement learning to jointly enhance legal knowledge grounding, task alignment, and reasoning capability. We construct a large-scale, high-quality legal corpus and design a structured post-training pipeline, enabling the model to effectively learn domain-specific knowledge and adapt to diverse legal tasks. We evaluate PoliLegalLM on three representative benchmarks, including LawBench, LexEval, and a real-world dataset, PoliLegal. Experimental results demonstrate that PoliLegalLM achieves strong and consistent performance, outperforming competitive models of similar scale and remaining highly competitive with significantly larger models, while achieving the best results on real-world legal scenarios. These results highlight the effectiveness of our training paradigm and the practical value of domain-specific LLMs for real-world legal applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。