arXiv:2507.05418cs.CLcs.AI2025-07被引 24

让大模型在多语言中精准推理,避免默认用英语思考。

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

  • 用多尺度对齐+翻译后语言一致性奖励,强制模型在目标语言推理。
  • 在地理事实问答任务中,模型在斯瓦希里语等低资源语言上准确率提升显著。
  • 新基准GeoFact-X支持五种语言的推理轨迹,适合评估跨语言推理能力。

大型语言模型在数学、事实问答和代码生成等领域表现强劲,但在不同语言上的推理能力仍较薄弱,尤其对斯瓦希里语、泰语等低资源语言,常误解提示或默认使用英语推理。这种对高资源语言的隐式偏倚损害了事实准确性、可解释性和可信度。我们提出M2A方法,结合多尺度多语言对齐与机器翻译问题的语言一致性奖励,训练模型直接且准确地在目标语言中推理。此外,现有评测仅关注最终答案,忽略推理是否在目标语言中发生。为此,我们构建了基于地理的多语言事实推理基准GeoFact-X,包含英文、印地语、日语、斯瓦希里语和泰语共五种语言的推理轨迹。实验表明,M2A显著提升了数学与事实推理任务中的多语言推理保真度,凸显推理感知的多语言强化学习对鲁棒跨语言泛化的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual question answering, and code generation, yet their ability to reason on these tasks in different languages remains underdeveloped. Especially for low-resource languages such as Swahili or Thai, LLMs can often misinterpret prompts or default to reasoning in English. This implicit bias toward high-resource languages undermines factual accuracy, interpretability, and trust. We propose M2A, a novel method that combines multi-scale multilingual alignment with language-consistency rewards on machine-translated questions, training models to reason directly and accurately in the target language. Furthermore, existing multilingual benchmarks only evaluate on final answers, overlooking whether reasoning occurs in the intended language. To close this gap, we introduce GeoFact-X, a geography-based multilingual factual reasoning benchmark together with reasoning traces in five languages: English, Hindi, Japanese, Swahili, and Thai. Our results show that M2A significantly enhances multilingual reasoning fidelity in both mathematical and factual reasoning tasks, highlighting that reasoning-aware multilingual reinforcement learning is crucial for robust cross-lingual generalization. https://jd730.github.io/projects/M2A_GeoFact-X

多语言推理强化学习低资源语言事实验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。