多语言微调能显著提升模型跨语言能力,即使只加一种非英语也有效。
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training

- 在220次实验中系统测试不同语言覆盖对模型的影响
- 加入非英语语言可提升英语表现和跨语言泛化,低资源语言受益最大
- 少量多语言数据即可带来显著效果,英语独占训练已不最优
尽管大语言模型广泛部署于多语言场景,其后训练流程仍以英语为主,导致语言间性能差异。我们基于220次监督微调实验,研究了训练语言覆盖范围、模型规模与任务领域之间的相互作用,使用涵盖数学推理和API调用任务的平行翻译多语言数据集,模型规模达80亿参数。结果表明,增加后训练中的语言覆盖在各类任务和模型规模下均有益,低资源语言受益最显著,高资源语言虽趋于饱和但未恶化。即使仅引入一种非英语语言,也能同时提升英语性能与跨语言泛化能力,说明纯英语后训练已非最优。此外,在足够多样性的语言背景下,零样本跨语言迁移可媲美甚至超越低多样性设置下的直接语言包含,但形态差异大且资源少的语言仍受限。
原文摘要 · Abstract (English)
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities across languages. We present a systematic, controlled study of the interplay between training language coverage, model scale, and task domain, based on 220 supervised fine-tuning runs on parallel translated multilingual data mixtures spanning mathematical reasoning and API calling tasks, with models up to 8B parameters. We find that increasing language coverage during post-training is largely beneficial across tasks and model scales, with low-resource languages benefiting the most and high-resource languages plateauing rather than degrading. Even minimal multilinguality helps: incorporating a single non-English language improves both English performance and cross-lingual generalization, making English-only post-training largely suboptimal. Moreover, at sufficient language diversity, zero-shot cross-lingual transfer can match or exceed the effects of direct language inclusion in a low-diversity setting, although gains remain limited for typologically distant, low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。