用翻译增强强化学习,提升多语言数学推理能力
TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning
- 以英语为桥梁,分步优化理解与推理过程
- 在多语言数学题上超越基线,跨语言泛化能力强
- 适合需要多语言推理的模型开发者与研究者
大型语言模型在英文数学推理中表现优异,但在多语言场景下仍存在显著性能差距,主要源于语言理解不足。为此,我们提出翻译增强策略优化(TAPO),一种基于GRPO的新型强化学习框架。TAPO采用显式对齐策略,以英语为枢纽,遵循‘先理解再推理’范式。关键在于引入逐步相对优势机制,分离理解与推理阶段,使翻译质量奖励可融入而不引发优化冲突。大量实验表明,TAPO有效融合语言理解与推理能力,兼容多种模型,在多语言数学推理和翻译任务中均优于基线方法,且在未见语言和跨领域任务中具有良好泛化性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable proficiency in English mathematical reasoning, yet a significant performance disparity persists in multilingual contexts, largely attributed to deficiencies in language understanding. To bridge this gap, we introduce Translation-Augmented Policy Optimization (TAPO), a novel reinforcement learning framework built upon GRPO. TAPO enforces an explicit alignment strategy where the model leverages English as a pivot and follows an understand-then-reason paradigm. Crucially, we employ a step-level relative advantage mechanism that decouples understanding from reasoning, allowing the integration of translation quality rewards without introducing optimization conflicts. Extensive experiments reveal that TAPO effectively synergizes language understanding with reasoning capabilities and is compatible with various models. It outperforms baseline methods in both multilingual mathematical reasoning and translation tasks, while generalizing well to unseen languages and out-of-domain tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。