开源中文法律大模型LuWen,提升法律推理与文本生成能力
WisdomInterrogatory (LuWen): An Open-Source Legal Large Language Model Technical Report

- 基于百川模型持续预训练+指令微调+检索增强生成
- 在5项法律任务中超越多个基线模型
- 适合法律AI研究者与司法智能化应用开发者
大语言模型在自然语言处理任务中表现出色,但在法律领域仍面临专业术语复杂、推理要求高、法律知识更新快等挑战。本文介绍开源中文法律大模型WisdomInterrogatory(LuWen),基于百川基础模型,通过大规模法律语料持续预训练、精心筛选的法律指令数据监督微调,以及融合全面法律知识库的检索增强生成技术构建。我们在五项代表性法律任务上评估了LuWen,涵盖法律判决预测、司法考试、法律文本摘要、法条问答和司法决策推理。实验结果表明,LuWen在各项任务中均优于多个强基线模型,验证了该方法在将通用语言模型适配法律领域的有效性。
原文摘要 · Abstract (English)
Large language models have demonstrated remarkable capabilities across a wide range of natural language processing tasks, yet their application in the legal domain remains challenging due to the specialized terminology, complex reasoning requirements, and rapidly evolving legal knowledge involved. In this paper, we present WisdomInterrogatory (LuWen), an open-source Chinese legal language model built upon the Baichuan foundation model through three key techniques: continual pre-training on a large-scale legal corpus, supervised fine-tuning with carefully curated legal instruction data, and retrieval-augmented generation integrated with a comprehensive legal knowledge base. We evaluate LuWen on five representative legal tasks spanning both prediction and generation settings, including legal judgment prediction, judicial examination, legal text summarization, law article question answering, and judicial decision reasoning. Experimental results show that LuWen outperforms several strong baselines, demonstrating the effectiveness of our approach in adapting general-purpose language models to the legal domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。