arXiv:2605.15613cs.CL2026-05

大模型偏爱英语,本土化训练比持续预训练更有效

Toward LLMs Beyond English-Centric Development

  • 分析开源大模型生成文本,发现严重英语偏好
  • 持续预训练在目标语言上不比从头训练省成本
  • 未来需为每种语言投入专门资源,而非依赖英语数据

通过对开源大语言模型生成文本序列的分析,我们发现大模型存在严重的英语中心偏差。尽管持续预训练常被用于将大模型适配到目标语言,但我们的研究显示,即使在提升目标语言文化理解方面,其成本优势也不明显,甚至不如从头训练。这一结果表明,未来大模型发展可能需要更多针对特定语言的专项投入,而非单纯依赖英语资源的扩展。

原文摘要 · Abstract (English)

Through an analysis of sequences generated by open-weight large language models (LLMs), we demonstrate that LLMs are heavily biased toward English. While continual pre-training is commonly used to adapt LLMs to a target language, we show that it does not offer a cost advantage over training from scratch, even for improving cultural understanding in the target language. These findings suggest that dedicated per-language investment may become increasingly important for future LLM development, rather than relying primarily on the expansion of English-centric resources.

大模型多语言偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。