arXiv:2604.17633cs.CL2026-04

揭示多语言预训练中翻译能力如何从抄写起步逐步演化

Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining

  • 通过细粒度训练监控,发现模型早期并行掌握语言基础与字面复制能力
  • 翻译能力分两阶段:初期依赖抄写和表面相似性,后期发展出更通用的翻译机制
  • 提出新数据集与分析方法,适合研究多语言模型学习动态的学者

大型语言模型展现出强大的跨语言能力。然而,以往研究多聚焦孤立因素或训练中的稀疏时间点,难以理解跨语言泛化如何形成,尤其在学习初期。为此,我们在九种不同语言上对一个17亿参数的多语言模型进行预训练,并以更高分辨率捕捉训练过程中的检查点。以词级翻译为测试基准,构建新数据集,通过行为分析、模型组件分析和参数消融实验,追踪翻译能力随训练演进的过程。结果发现,模型在早期快速获得基本语言能力,同时具备逐标记复制能力;翻译能力发展分为两个阶段:初始阶段以复制和表面相似性主导,第二阶段则发展出更泛化的翻译机制,同时复制能力被精细化。这些发现提供了多语言预训练中跨语言泛化演化的细粒度视角。

原文摘要 · Abstract (English)

Large language models exhibit impressive cross-lingual capabilities. However, prior work analyzes this phenomenon through isolated factors and at sparse points during training, limiting our understanding of how cross-lingual generalization emerges--particularly in the early phases of learning. To study the early trajectory of linguistic and translation capabilities, we pretrain a multilingual 1.7B model on nine diverse languages, capturing checkpoints at a much finer granularity. We use word-level translation as a testbed, introducing a novel dataset to trace how translation develops over training through behavioral analyses, model-component analysis, and parameter-based ablations. We find that the model quickly acquires basic linguistic capabilities in parallel with token-level copying, while translation develops in two distinct phases: an initial phase dominated by copying and surface-level similarities, and a second phase in which more generalizing translation mechanisms are developed while copying is refined. Together, these findings provide a fine-grained view of how cross-lingual generalization develops during multilingual pretraining.

多语言模型翻译机制预训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。