arXiv:2602.21887cs.CL2026-02被引 3

让大模型在推理时自主选语言,提升思考效率与多语种表现。

ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection

  • 训练中动态选择思考语言,增强强化学习探索能力。
  • 相同预算下超越纯英文训练,多语言表现更稳定。
  • 适合需要多语种推理能力的研究与应用

当前大型推理模型(LRMs)在强化学习(RL)后训练后展现出强大解决复杂任务的能力。然而,以往工作主要聚焦于英语推理,期望获得最优性能,尽管多语言思考已显示出潜在优势,且全球用户对母语思考记录有实际需求。本文提出ExpLang,一种新型的大模型后训练流程,通过在强化学习过程中实现策略内思考语言选择,以提升探索与利用效果。实验表明,在相同训练预算下,该方法持续优于仅用英语训练的模型,且对已见和未见语言均保持高语言合规性。分析显示,将思考语言选择作为强化学习中的动作,有效扩展了探索空间,增强了非英语语言的优势,从而提升利用效果。该方法与多数强化学习算法正交,为利用多语言性提升大型推理模型开辟了新视角。

原文摘要 · Abstract (English)

Current large reasoning models (LRMs) have shown strong ability on challenging tasks after reinforcement learning (RL) based post-training. However, previous work mainly focuses on English reasoning in expectation of the strongest performance, despite the demonstrated potential advantage of multilingual thinking, as well as the requirement for native thinking traces by global users. In this paper, we propose ExpLang, a novel LLM post-training pipeline that enables on-policy thinking language selection to improve exploration and exploitation during RL with the use of multiple languages. The results show that our method steadily outperforms English-only training with the same training budget, while showing high thinking language compliance for both seen and unseen languages. Analysis shows that, by enabling on-policy thinking language selection as an action during RL, ExpLang effectively extends the RL exploration space with diversified language preference and improves the RL exploitation outcome with leveraged non-English advantage. The method is orthogonal to most RL algorithms and opens up a new perspective on using multilinguality to improve LRMs.

大模型推理多语言强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。