arXiv:2412.12956cs.CL2024-12中稿 · NoDaLiDa 2025

丹麦语大模型SnakModel通过海量语料训练,提升小语种NLP性能。

SnakModel: Lessons Learned from Training an Open Danish Large Language Model

  • 基于Llama2-7B架构,用136亿词丹文语料持续预训练。
  • 在8个文化相关任务中表现最优,优于多个同类模型。
  • 开源模型、语料与代码,助力小语种语言研究与发展。

我们提出SnakModel,一个基于Llama2-7B的丹麦语大语言模型,通过136亿词丹麦语文本进行持续预训练,并在370万条丹麦语指令上进一步微调。由于小语种社区尚无成熟的LLM构建规范,我们系统考察了建模与训练决策对下游性能的影响,涵盖:(1) 来自多样化来源的严格筛选丹麦语文本语料构建;(2) 语言建模与指令微调过程,包括中间训练动态分析及超参数消融实验;(3) 在八个语言与文化特异性任务上的评估。实验表明,SnakModel整体表现最佳,超越多个同类Llama2-7B模型。通过开放SnakModel、主要预训练语料及配套代码,我们希望推动丹麦语自然语言处理研究,并为资源有限语言提供可借鉴的训练指南。

原文摘要 · Abstract (English)

We present SnakModel, a Danish large language model (LLM) based on Llama2-7B, which we continuously pre-train on 13.6B Danish words, and further tune on 3.7M Danish instructions. As best practices for creating LLMs for smaller language communities have yet to be established, we examine the effects of early modeling and training decisions on downstream performance throughout the entire training pipeline, including (1) the creation of a strictly curated corpus of Danish text from diverse sources; (2) the language modeling and instruction-tuning training process itself, including the analysis of intermediate training dynamics, and ablations across different hyperparameters; (3) an evaluation on eight language and culturally-specific tasks. Across these experiments SnakModel achieves the highest overall performance, outperforming multiple contemporary Llama2-7B-based models. By making SnakModel, the majority of our pre-training corpus, and the associated code available under open licenses, we hope to foster further research and development in Danish Natural Language Processing, and establish training guidelines for languages with similar resource constraints.

大模型小语种语言建模丹麦语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。