arXiv:2504.11833cs.CL2025-04被引 15

用多语言推理能让大模型表现更好,比英语还强近10分。

Could Thinking Multilingually Empower LLM Reasoning?

  • 用多种语言做推理任务,突破英语单一语言上限。
  • 多语言推理在不同翻译质量下仍表现稳定,最高提升近10分。
  • 现有答案选择方法无法达到此上限,需新思路。

以往研究指出大语言模型存在显著的‘英语偏好’,即任务用英语呈现时表现更优。然而,我们观察到某些其他语言在推理任务中反而能带来比英语更好的效果。这一现象尚未深入探索。本文研究多语言推理的上限,发现其性能可显著(接近10 Acc@$k$点)且稳健地超越纯英语推理,对翻译质量和语言选择变化具有较强容忍度。同时分析了达到该上限的机制与挑战,并发现现有答案选择方法因自身局限和偏差,无法触及这一上限。这些发现为未来全面挖掘大模型多语言推理潜力提供了重要方向。

原文摘要 · Abstract (English)

Previous work indicates that large language models exhibit a significant "English bias", i.e. they often perform better when tasks are presented in English. Interestingly, we have observed that using certain other languages in reasoning tasks can yield better performance than English. However, this phenomenon remains under-explored. In this paper, we explore the upper bound of harnessing multilingualism in reasoning tasks, suggesting that multilingual reasoning promises significantly (by nearly 10 Acc@$k$ points) and robustly (tolerance for variations in translation quality and language choice) higher upper bounds than English-only reasoning. Besides analyzing the reason behind the upper bound and challenges in reaching it, we also find that common answer selection methods cannot achieve this upper bound, due to their limitations and biases. These insights could pave the way for future research aimed at fully harnessing the potential of multilingual reasoning in LLMs.

多语言推理大模型性能提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。