arXiv:2509.15811cs.CLcs.AI2025-09Conference of the …被引 2

跨语言奖励模型提升数学推理能力,尤其在低采样预算下增强英文表现。

Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning

  • 构建跨语言奖励模型,统一评估多语言生成答案质量。
  • 相比单语言奖励模型,数学推理性能显著提升,英语言在低预算时受益最大。
  • 利用不同语言的互补推理优势,为多语言模型优化提供新思路。

尽管大语言模型(LLMs)的推理能力持续进步,但其在多语言模型中跨语言的表现差异尚不明确,且不同语言是否产生互补性推理路径仍不清楚。为此,我们训练了一个跨语言奖励模型,用于对多语言生成的回答进行排序。结果表明,与单一语言奖励建模相比,该模型显著提升了数学推理性能,甚至惠及高资源语言。虽然英语在多语言模型中通常表现最优,但在低采样预算下,跨语言采样对英语的增益尤为明显。研究揭示了通过利用多种语言的互补优势来提升多语言推理的新机遇。

原文摘要 · Abstract (English)

While the reasoning abilities of large language models (LLMs) continue to advance, it remains unclear how such ability varies across languages in multilingual LLMs and whether different languages produce reasoning paths that complement each other. To investigate this question, we train a reward model to rank generated responses for a given question across languages. Our results show that our cross-lingual reward model substantially improves mathematical reasoning performance compared to using reward modeling within a single language, benefiting even high-resource languages. While English often exhibits the highest performance in multilingual models, we find that cross-lingual sampling particularly benefits English under low sampling budgets. Our findings reveal new opportunities to improve multilingual reasoning by leveraging the complementary strengths of diverse languages.

多语言数学推理奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。