arXiv:2603.01691cs.CLcs.LG2026-03被引 2

为斯洛文尼亚语打造高性能开源大模型,效果媲美大厂闭源模型。

Building a Strong Instruction Language Model for a Less-Resourced Language

  • 三阶段持续预训练+两阶段有监督微调,适配低资源语言
  • 120亿参数模型在三项评测中超越同规模Gemma 3,胜率超60%
  • 适合研究低资源语言AI、本地化大模型部署的团队使用

大型语言模型已成为自然语言处理和人工智能的核心工具,但当前开源模型主要基于英语文本训练,导致对低资源语言性能较差。本文提出一套方法论,成功将大模型适配至斯洛文尼亚语,并以该语言为例进行验证。我们构建了名为GaMS3-12B的生成式模型,参数量达120亿,是当前同参数规模下表现最优的开源斯洛文尼亚语模型。通过三阶段持续预训练(基于Gemma 3)及两阶段有监督微调,模型在包含1400亿词符的斯洛文尼亚语、英语、波斯尼亚语、塞尔维亚语和克罗地亚语文本上进行训练,并使用超过20万条英斯双语指令数据进行微调。在Slovenian-LLM-Eval、英译斯洛文尼亚语及斯洛文尼亚语LLM竞技场三项评测中,该模型均优于12B Gemma 3,在斯洛文尼亚语竞技场中表现接近商业级GPT-4o,胜率超过60%。

原文摘要 · Abstract (English)

Large language models (LLMs) have become an essential tool for natural language processing and artificial intelligence in general. Current open-source models are primarily trained on English texts, resulting in poorer performance on less-resourced languages and cultures. We present a set of methodological approaches necessary for the successful adaptation of an LLM to a less-resourced language, and demonstrate them using the Slovene language. We present GaMS3-12B, a generative model for Slovene with 12 billion parameters, and demonstrate that it is the best-performing open-source model for Slovene within its parameter range. We adapted the model to the Slovene language using three-stage continual pre-training of the Gemma 3 model, followed by two-stage supervised fine-tuning (SFT). We trained the model on a combination of 140B Slovene, English, Bosnian, Serbian, and Croatian pretraining tokens, and over 200 thousand English and Slovene SFT examples. We evaluate GaMS3-12B on the Slovenian-LLM-Eval datasets, English-to-Slovene translation, and the Slovene LLM arena. We show that the described model outperforms 12B Gemma 3 across all three scenarios and performs comparably to much larger commercial GPT-4o in the Slovene LLM arena, achieving a win rate of over 60 %.

低资源语言大模型适配斯洛文尼亚语开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。