提出多语言数学推理评估新方法,提升评测鲁棒性。
MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
- 为每道题生成五种数字/名称/上下文变体,增强评估多样性。
- 低资源语言在不同数字变体下性能下降超30%,体现评测偏差。
- 开源模型如GPT-OSS和DeepSeek表现更稳健,适合多语言研究。
大语言模型在数学推理方面取得显著进展,但多语言评估基准在难度与更新速度上仍落后于英语。近期GSM-Symbolic揭示了同一问题不同实例化版本间模型表现存在显著差异,但仅限英文测试。本文提出MGSM-Pro,基于MGSM数据集扩展GSM-Symbolic方法,为每道题生成五种变体(改变名称、数字及无关上下文)。跨九种语言的评估显示,许多低资源语言在不同数字变体下性能大幅下降;高资源语言下的鲁棒性不保证低资源语言同样适用。此外,闭源模型如Gemini 2.5 Flash与GPT-4.1对数字变化敏感,而Gemini 3.0 Pro更具鲁棒性;开源模型中,GPT-OSS 120B与DeepSeek v3表现更优。建议至少使用五种数字变体进行评估,以获得更可靠、真实的数学推理能力衡量。
原文摘要 · Abstract (English)
Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set. We further find that models robustness in HRL setting do not necessarily translate to LRL. Moreover, proprietary models, such as Gemini 2.5 Flash and GPT-4.1 are less robust to digit, whereas Gemini 3.0 Pro is more robust. Among open models, GPT-OSS 120B and DeepSeek v3 show stronger robustness. Based on these findings, we recommend evaluating each problem using at least five digit-varying instantiations to obtain a more robust and realistic assessment of math reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。