arXiv:2505.11810cs.CL2025-05被引 3

用18亿参数打造高效古典中文大模型,性能超通用模型。

Efficiently Building a Domain-Specific Large Language Model from Scratch: A Case Study of a Classical Chinese Large Language Model

  • 从零构建专用于古典中文的模型,优化设计与训练流程。
  • 在断句、典故识别等任务上接近或超越人类水平。
  • 适合古籍整理、辞书编纂等传统文化研究场景。

通用大语言模型在自然语言处理任务中表现优异,甚至在某些任务上达到或超过人类水平。然而,当应用于古典中文等特定领域时,其效果往往不理想;微调开源基础模型也难以充分融入领域知识。为此,本研究开发了专用于古典中文理解与生成的大语言模型AI Taiyan。实验表明,在合理设计模型结构、数据处理、预训练与微调的基础上,仅需18亿参数即可取得良好效果。在古典中文语言处理的关键任务如断句、典故识别、词义解释及古今汉语互译中,该模型显著优于通用大模型和传统领域模型,性能接近或超过人类基准。本研究为高效构建专用领域大模型提供了参考。此外,论文结合案例探讨了该模型在古籍校勘、辞书编辑与语言研究中的应用价值。

原文摘要 · Abstract (English)

General-purpose large language models demonstrate notable capabilities in language comprehension and generation, achieving results that are comparable to, or even surpass, human performance in many natural language processing tasks. Nevertheless, when general models are applied to some specific domains, e.g., Classical Chinese texts, their effectiveness is often unsatisfactory, and fine-tuning open-source foundational models similarly struggles to adequately incorporate domain-specific knowledge. To address this challenge, this study developed a large language model, AI Taiyan, specifically designed for understanding and generating Classical Chinese. Experiments show that with a reasonable model design, data processing, foundational training, and fine-tuning, satisfactory results can be achieved with only 1.8 billion parameters. In key tasks related to language processing of Classical Chinese such as punctuation, identification of allusions, explanation of word meanings, and translation between ancient and modern Chinese, this model exhibits a clear advantage over both general-purpose large models and domain-specific traditional models, achieving levels close to or surpassing human baselines. This research provides a reference for the efficient construction of specialized domain-specific large language models. Furthermore, the paper discusses the application of this model in fields such as the collation of ancient texts, dictionary editing, and language research, combined with case studies.

古典中文大模型知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。