通过数学等价变换测试大模型推理鲁棒性,发现其易受语言和参数扰动影响。
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 设计数学等价但语言参数不同的题目,评估大模型对非数学扰动的敏感度。
- 18个模型在变体题上表现下降,旗舰模型O3准确率降幅达12.9个百分点。
- 新基准数据集PutnamGAP可揭示模型弱点,助力提升数学推理能力。
本文提出一种超越传统方法的系统性框架,通过在具有语言和参数变化但数学等价的高阶数学问题上对大语言模型进行压力测试,评估其数学推理鲁棒性。该方法能有效衡量模型对非数学扰动的敏感度,从而更准确地评估其真实推理能力。基于此方法,我们构建了名为PutnamGAP的新基准数据集,包含竞赛级数学问题的多种数学等价变体。在该数据集上,我们评估了18个代表性商业与开源模型的性能。结果表明,所有模型在变体题上均出现显著性能下降:OpenAI旗舰推理模型O3在原题上得分为51.5%,在表面重命名变体上下降4.7个百分点,在参数变体上下降12.9个百分点;较小模型表现更差。整体结果验证了新评估方法的有效性,有助于深化对大模型数学推理鲁棒性的理解,并为后续改进提供新思路。
原文摘要 · Abstract (English)
In this paper, we introduce a systematic framework beyond conventional method to assess LLMs' mathematical-reasoning robustness by stress-testing them on advanced math problems that are mathematically equivalent but with linguistic and parametric variation. These transformations allow us to measure the sensitivity of LLMs to non-mathematical perturbations, thereby enabling a more accurate evaluation of their mathematical reasoning capabilities. Using this new evaluation methodology, we created PutnamGAP, a new benchmark dataset with multiple mathematically-equivalent variations of competition-level math problems. With the new dataset, we evaluate multiple families of representative LLMs and examine their robustness. Across 18 commercial and open-source models we observe sharp performance degradation on the variants. OpenAI's flagship reasoning model, O3, scores 51.5% on the originals but drops by 4.7 percentage points on surface-renaming variants, and by 12.9 percentage points on parametric variants, while smaller models fare far worse. Overall, the results show that the proposed new evaluation methodology is effective for deepening our understanding of the robustness of LLMs and generating new insights for further improving their mathematical reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。