arXiv:2512.21580cs.CL2025-12被引 1

1.5B参数模型用2.5万亿词训练,多语言性能超越更大模型。

Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM

  • 分两阶段训练:先均衡多语言对齐,再强化英语提升跨语言能力。
  • 仅用2.5万亿词训练,性能超LLaMA3.2-1B(9T)和Qwen2.5-1.5B(18T)。
  • 俄语表现领先,1-2B级模型中在MERA基准上达当前最优。

我们提出Gamayun,一个从头训练的1.5B参数多语言语言模型,使用2.5万亿个标记进行训练。为适应资源受限环境,该模型采用创新的两阶段预训练策略:第一阶段为平衡多语言训练以实现跨语言对齐;第二阶段通过高质量英语数据增强,实现性能向其他语言迁移。模型支持12种语言,尤其侧重俄语。尽管训练预算远低于同类模型,其在所有评估基准上均优于LLaMA3.2-1B(9T tokens),并在广泛英语与多语言任务上超越Qwen2.5-1.5B(18T tokens)。在多数任务上达到或超过Qwen3(36T tokens)水平,尤其在俄语领域表现突出,在同规模模型中于MERA基准取得最佳成绩。

原文摘要 · Abstract (English)

We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small non-English-centric LLMs by adopting a novel two-stage pre-training strategy: balanced multilingual training for cross-lingual alignment, followed by high-quality English enrichment to transfer performance gains across languages. Our model supports 12 languages, with special focus on Russian. Despite a significantly smaller training budget than comparable models, Gamayun outperforms LLaMA3.2-1B (9T tokens) on all considered benchmarks, and surpasses Qwen2.5-1.5B (18T tokens) on a wide range of English and multilingual tasks. It matches or exceeds Qwen3 (36T tokens) on most tasks outside advanced STEM, achieving state-of-the-art results in Russian, including the MERA benchmark, among the models of comparable size (1-2B parameters).

多语言小模型效率俄语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。