arXiv:2602.05940cs.CL2026-02被引 1

让多语言模型真正用母语推理,准确率提升10.3个百分点。

R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning

  • 分离理解与推理优化,通过自动生成英文提示恢复目标语言强化信号。
  • 在五种语言上平均提升10.3%准确率,同时保持近乎完美的语言一致性。
  • 无需外部数据或模型,适用于多种大模型和跨领域任务。

大型推理模型在处理非英语问题时常默认使用英语推理,导致目标语言性能显著下降。即使使用相同语言推理,语义等价的中英文问题仍存在明显准确率差距。这揭示了两个瓶颈:目标语言理解与目标语言推理。现有方法通常仅优化其中之一,简单组合效果有限,因答案正确性无法区分理解失败与推理失败。我们提出R3S,一种解耦两能力优化的强化学习框架。R3S通过英语可解性过滤精炼翻译奖励,并利用自生成英文提示恢复目标语言的强化信号。该方法无需外部模型反馈或多语言训练数据。在三种主干模型和五种语言上的实验表明,R3S在MMATH上相较目标语言强化信号基线平均提升10.3个百分点,同时保持近似完美语言一致性。在MMLU-ProX上的持续增益进一步验证其在数学以外任务的泛化能力。

原文摘要 · Abstract (English)

Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reasoning in the question language. Even with the same reasoning language, semantically equivalent English and non-English questions still exhibit a clear performance gap. Together, these phenomena reveal two distinct bottlenecks: target-language question understanding and target-language reasoning. Existing methods typically optimize only one of these capabilities. However, simply combining them may not be sufficient to optimize both effectively, as answer correctness alone cannot distinguish failures in question understanding from those in reasoning. We propose R3S, a reinforcement learning framework that disentangles the optimization of the two capabilities. R3S refines translation rewards derived from downstream reasoning accuracy through English-solvability filtering and recovers target-language RLVR signals using self-generated English hints. Together, these designs require neither external model feedback nor external multilingual training data. Experiments across three backbone models and five languages show that R3S improves language-consistent accuracy over the target-language RLVR baseline on MMATH by an average of 10.3 percentage points, while maintaining near-perfect language consistency. Consistent gains on MMLU-ProX further demonstrate its generalization beyond math problems.

多语言推理强化学习语言一致性模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。