arXiv:2512.10772cs.CLcs.AI2025-12被引 3

通过扩大模型规模,更高效地适配低资源语言。

Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation

  • 用更大模型放大英文基础模型,替代传统持续预训练。
  • 大模型在少量目标语料下表现超小模型,且不丢英语能力。
  • 可合并多个语言模型构建灵活多语言系统,仍有优化空间。

实现高性能的多语言模型,尤其是中低资源语言,仍是挑战。当前大规模多语言模型性能仍低于针对特定语言的微调模型,尤其在小模型规模下。本文研究将模型规模扩展作为高效适配新语言的策略。通过约等计算量(FLOP-matched)的缩放实验,我们验证:在接触足够目标语言数据后,扩大的英文基座模型能媲美甚至超越持续预训练大量数据的小模型,证明了缩放对数据效率的优势。缩放还能更好保留英语能力,减轻灾难性遗忘。最后,我们探索将此类扩展后的语言专用模型进行合并,构建模块化多语言系统。发现合并效果虽不及联合训练,但扩大后的合并优于小模型合并。不同合并方法间性能差异显著,提示可通过专为语言级融合设计的方法进一步提升。

原文摘要 · Abstract (English)

Achieving high-performing language models which include medium- and lower-resource languages remains a challenge. Massively multilingual models still underperform compared to language-specific adaptations, especially at smaller model scales. In this work, we investigate scaling as an efficient strategy for adapting pretrained models to new target languages. Through comprehensive scaling ablations with approximately FLOP-matched models, we test whether upscaling an English base model enables more effective and resource-efficient adaptation than standard continued pretraining. We find that, once exposed to sufficient target-language data, larger upscaled models can match or surpass the performance of smaller models continually pretrained on much more data, demonstrating the benefits of scaling for data efficiency. Scaling also helps preserve the base model's capabilities in English, thus reducing catastrophic forgetting. Finally, we explore whether such scaled, language-specific models can be merged to construct modular and flexible multilingual systems. We find that while merging remains less effective than joint multilingual training, upscaled merges perform better than smaller ones. We observe large performance differences across merging methods, suggesting potential for improvement through merging approaches specialized for language-level integration.

语言适配模型缩放多语言融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。