多语言数学推理评测基准,揭示大模型在18种语言中的表现差异。
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
- 构建覆盖18语言、4难度等级的多语言数学推理测试集
- 顶尖模型最高仅40%准确率,高难度下性能显著下降
- 发现语言一致性差、思维长度受语言影响,适合多语言研究者
本文提出PolyMath,一个涵盖18种语言和4个从易到难难度级别的多语言数学推理基准。该基准确保难度全面性、语言多样性与高质量翻译,成为当前大模型推理能力评估的重要工具。我们对先进大模型进行了全面评估,结果显示即使是最先进的Qwen-3-235B-A22B-Thinking和Gemini-2.5-pro,得分也仅分别为54.6和52.2,且在最高等级下准确率约为40%。从语言角度看,研究揭示了当前大模型在多语言推理中的若干关键挑战:(1)不同语言间推理表现差异显著;(2)输入输出语言一致性较低,且可能与性能相关;(3)不同语言下思维链长度存在明显差异。此外,我们发现控制指令中的输出语言可能影响推理表现,尤其在低资源语言中效果显著,提示了一条提升大模型多语言能力的可行路径。
原文摘要 · Abstract (English)
In this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs. We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning: (1) Reasoning performance varies widely across languages for current LLMs; (2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance; (3) The thinking length differs significantly by language for current LLMs. Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。