arXiv:2509.23657cs.CL2025-09被引 11

强化学习让大模型跨语言推理更准更强

Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs

  • 用强化学习替代监督微调,提升跨语言推理能力
  • 非英语数据训练的RL模型整体表现更好,跨语言泛化更强
  • 适合关注多语言公平性与通用推理能力的研究者

提升大语言模型(LLMs)的复杂推理能力备受关注。尽管强化学习(RL)在改善复杂推理方面表现优异,但其在跨语言泛化方面相较于监督微调(SFT)的效果尚不明确。本文首次系统研究了RL与SFT在跨语言推理泛化上的差异。以Qwen2.5-3B-Base为基础模型,在涵盖数学推理、常识推理和科学推理的多语言基准上进行实验。结果发现:(1) 使用RL不仅准确率更高,且跨语言泛化能力显著优于SFT;(2) 在非英语数据上训练的RL模型整体性能及泛化能力优于英语数据训练的RL模型,而SFT则无此现象。通过深入机制分析,揭示了RL在跨语言场景中具备更鲁棒的推理策略。研究为实现更公平高效的多语言推理提供了关键指导。

原文摘要 · Abstract (English)

Enhancing the complex reasoning capabilities of Large Language Models (LLMs) attracts widespread attention. While reinforcement learning (RL) has shown superior performance for improving complex reasoning, its impact on cross-lingual generalization compared to Supervised Fine-Tuning (SFT) remains unexplored. We present the first systematic investigation into cross-lingual reasoning generalization of RL and SFT. Using Qwen2.5-3B-Base as our foundation model, we conduct experiments on diverse multilingual reasoning benchmarks, including math reasoning, commonsense reasoning, and scientific reasoning. Our investigation yields two significant findings: (1) Tuning with RL not only achieves higher accuracy but also demonstrates substantially stronger cross-lingual generalization capabilities compared to SFT. (2) RL training on non-English data yields better overall performance and generalization than training on English data, which is not observed with SFT. Furthermore, through comprehensive mechanistic analyses, we explore the underlying factors of RL's superiority and generalization across languages. Our results provide compelling evidence that RL enables the model with more robust reasoning strategies, offering crucial guidance for more equitable and effective multilingual reasoning.

强化学习跨语言推理能力多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。