arXiv:2412.06484cs.CL2024-12被引 13

为挪威语等小语种训练大模型,提升效果与效率

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

  • 分三阶段持续训练,优化小语种模型性能
  • 推出114亿参数的NorMistral-11B模型,支持三种挪威语言
  • 开源模型,助力低资源语言研究与应用

训练大型语言模型需要海量数据,这对挪威语等使用人数较少的语言构成挑战,而更罕见的低资源语言如北萨米语则更为困难。为此,我们提出一种新颖的三阶段持续训练方法,显著提升了目标语言的下游性能和推理效率。基于研究发现,我们训练、评估并公开发布了针对挪威语书面语(Bokmål)、新挪威语(Nynorsk)和北萨米语的生成式语言模型:拥有114亿参数的NorMistral-11B。

原文摘要 · Abstract (English)

Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern Sámi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokmål, Nynorsk, and Northern Sámi with 11.4 billion parameters: NorMistral-11B.

语言模型小语种持续训练开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。