通过自反馈机制提升多语言数学推理一致性。
Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning
- 以主语言为枢纽,翻译问题并监督推理过程。
- 在多个低资源语言上提升数学推理准确率。
- 无需外部答案或奖励模型,适合多语言场景。
尽管大型语言模型展现了强大的推理能力,但实证表明其并非如预期般具备语言无关性,在多语言环境下表现下降,尤其在低资源语言上更为明显。我们归因于模型在多语言理解与推理对齐上的不一致。为此,提出枢纽对齐自反馈多语言推理方法(PASMR),将模型主语言设为枢纽语言。训练时,先将问题翻译至枢纽语言,以促进推理模式对齐;随后用枢纽语言的推理答案监督目标语言的推理过程,建立跨语言自反馈机制,无需依赖外部正确答案或奖励模型。大量实验表明,该方法显著提升了模型对问题的理解与推理能力,带来明显任务性能提升。
原文摘要 · Abstract (English)
Despite the impressive reasoning abilities demonstrated by large language models (LLMs), empirical evidence indicates that they are not language agnostic as expected, leading to performance declines in multilingual settings, especially for low-resource languages. We attribute the decline to the model's inconsistent multilingual understanding and reasoning alignment. To address this, we present Pivot-Aligned Self-Feedback Multilingual Reasoning (PASMR), aiming to improve the alignment of multilingual math reasoning abilities in LLMs. This approach designates the model's primary language as the pivot language. During training, the model first translates questions into the pivot language to facilitate better alignment of reasoning patterns. The reasoning process in the target language is then supervised by the pivot language's reasoning answers, thereby establishing a cross-lingual self-feedback mechanism without relying on external correct answers or reward models. Extensive experimental results demonstrate that our method enhances both the model's understanding of questions and its reasoning capabilities, leading to notable task improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。