首个专用于生成中文裁判文书的轻量级大模型,提升法律文本自动化质量
ShiZhi: A Chinese Lightweight Large Language Model for Court View Generation
- 基于11万+真实案件构建中文裁判文书数据集
- 轻量模型在生成任务中达70.00 ROUGE-1,定罪预测准确率86.48%
- 小模型也能生成合乎法律逻辑的裁判意见,适合司法AI落地
刑事裁判文书生成(CVG)是法律人工智能的基础任务,旨在自动生成法律文书中的“法院观点”部分。由于案件事实多样复杂,直接从原始事实生成常受限于表现。本文提出首个专为裁判文书生成设计的中文大语言模型ShiZhi,构建了包含超过11万例案件的中文裁判文书生成数据集CCVG,每例包含事实描述与对应裁判观点。基于该数据集,ShiZhi在裁判文书生成任务上达到70.00 ROUGE-1和67.85 BLEU-1,在定罪预测任务上实现86.48%准确率与92.75%宏平均F1值。实验表明,仅用高质量领域数据训练的小型模型亦可生成合理且符合法律逻辑的裁判观点。模型与数据集已开源。
原文摘要 · Abstract (English)
Criminal Court View Generation (CVG) is a fundamental task in legal artificial intelligence, aiming to automatically generate the "Court View" section of a legal case document. Generating court views is challenging due to the diversity and complexity of case facts, and directly generating from raw facts may limit performance. In this paper, we present ShiZhi, the first large language model (LLM) specifically designed for court view generation. We construct a Chinese Court View Generation dataset, CCVG, of more than 110K cases, each containing fact descriptions paired with corresponding court views. Based on this dataset, ShiZhi achieving 70.00 ROUGE-1 and 67.85 BLEU-1 on court view generation, as well as 86.48\% accuracy with 92.75\% macro F1 on charge prediction. Experimental results demonstrate that even a small LLM can generate reasonable and legally coherent court views when trained on high-quality domain-specific data. Our model and dataset are available at \href{https://github.com/ZhitianHou/ShiZhi}{https://github.com/ZhitianHou/ShiZhi}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。