arXiv:2608.13698cs.CLcs.LG2026-08

跨语言训练让大模型推理能力显著提升,但需警惕语言间性能退化。

GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

  • 在多语言环境下用GRPO优化模型推理能力
  • 母语训练仅留小幅差距,跨语言迁移效果显著
  • 需全面评估以发现特定语言的性能退化

基于可验证奖励的强化学习(RLVR)常通过分组相对策略优化(GRPO)来提升预训练语言模型的推理能力,但现有研究主要集中于英语。本文开展大规模实证研究,覆盖多种基础模型、训练语言及不同推理语言奖励,在多语言和非英语场景下评估GRPO表现。结果表明,以母语进行推理训练时,性能与英语训练仅存在微小差距;同时观察到强跨语言迁移能力:一种语言的训练往往能提升多种其他语言的表现。然而,具体趋势高度依赖模型和语言。某些情况下,特定语言训练反而导致其他语言能力严重退化。分析显示,非英语环境下的RLVR可带来广泛跨语言收益,但必须进行广泛评估以识别语言特异性退化问题。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.

强化学习多语言推理能力GRPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。