arXiv:2504.16677cs.CLcs.AI2025-04被引 7

揭秘大模型多语言训练中跨语言迁移的真正机制

A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics

  • 在真实训练环境下分析多语言数据对跨语言迁移的影响
  • 发现模型性能受任务类型与训练方式组合共同决定
  • 给出实际应用中提升跨语言能力的有效配置建议

为使大语言模型在全球范围内有用,它们通常在多语言数据上进行指令微调。尽管此类后训练普遍应用,但对促进跨语言迁移的动态机制仍缺乏清晰理解。本研究在真实后训练设置下考察跨语言迁移(CLT)动态。我们分析了两种规模达350亿参数的模型家族,在三种生成任务(摘要、指令遵循、数学推理)中,采用不同复杂度的多语言数据混合,并在单任务与多任务指令调优场景下进行实验。结果表明,跨语言迁移与多语言性能无法仅由单一变量解释,其表现取决于后训练设置的组合。最终,我们识别出实践中实现有效跨语言迁移的关键条件。

原文摘要 · Abstract (English)

In order for large language models to be useful across the globe, they are fine-tuned to follow instructions on multilingual data. Despite the ubiquity of such post-training, a clear understanding of the dynamics that enable cross-lingual transfer remains elusive. This study examines cross-lingual transfer (CLT) dynamics in realistic post-training settings. We study two model families of up to 35B parameters in size trained on carefully controlled mixtures of multilingual data on three generative tasks with varying levels of complexity (summarization, instruction following, and mathematical reasoning) in both single-task and multi-task instruction tuning settings. Overall, we find that the dynamics of cross-lingual transfer and multilingual performance cannot be explained by isolated variables, varying depending on the combination of post-training settings. Finally, we identify the conditions that lead to effective cross-lingual transfer in practice.

跨语言迁移指令微调多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。