专为文言文设计的大模型,性能远超现有水平。
WenyanGPT: A Large Language Model for Classical Chinese Tasks

- 在中文大模型基础上继续预训练并指令微调,专注文言文理解。
- 在自建评测集上显著优于当前先进模型,多项任务领先。
- 适合古籍研究、文化遗产数字化等领域的学者与开发者。
文言文作为中华文化的载体,在古代文献传承与研究中具有关键作用。然而,现有自然语言处理模型主要针对现代汉语优化,对文言文处理效果不佳。本文提出一套完整的文言文语言处理解决方案:基于LLaMA3-8B-Chinese模型进行持续预训练和指令微调,构建了专门用于文言文任务的大语言模型WenyanGPT。同时,我们开发了一个评估基准数据集WenyanBENCH。在该数据集上的实验结果表明,WenyanGPT在各类文言文任务中显著优于当前先进大模型。我们公开了模型的训练数据、指令微调数据及评估基准数据集,以推动文言文处理领域的进一步研究与发展。
原文摘要 · Abstract (English)
Classical Chinese, as the core carrier of Chinese culture, plays a crucial role in the inheritance and study of ancient literature. However, existing natural language processing models primarily optimize for Modern Chinese, resulting in inadequate performance on Classical Chinese. This paper presents a comprehensive solution for Classical Chinese language processing. By continuing pre-training and instruction fine-tuning on the LLaMA3-8B-Chinese model, we construct a large language model, WenyanGPT, which is specifically designed for Classical Chinese tasks. Additionally, we develop an evaluation benchmark dataset, WenyanBENCH. Experimental results on WenyanBENCH demonstrate that WenyanGPT significantly outperforms current advanced LLMs in various Classical Chinese tasks. We make the model's training data, instruction fine-tuning data\footnote, and evaluation benchmark dataset publicly available to promote further research and development in the field of Classical Chinese processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。