重新评估GSM-Symbolic测试,发现模型推理能力结论需谨慎,因统计和数据偏差影响结果。
The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

- 采用带随机效应的广义线性混合模型,更精准分析模型性能变化。
- 仅8个模型在原提示格式下表现显著下降,且多数差异由数值分布偏移导致。
- 揭示模型特异性失败模式,提醒避免对大模型推理能力的笼统判断。
GSM-Symbolic基准(Mirzadeh等,2025)在模板生成的GSM8K变体上观察到25个大型语言模型性能普遍下降,据此认为模型缺乏真实推理能力。本文指出该结论建立在薄弱的统计基础之上。通过使用带有逐题随机效应的自助法广义线性混合模型重新评估20个开源权重模型,我们发现仅8个模型在原始提示格式下表现出统计显著的性能变化。此外,我们识别出一个此前未被注意的因素:主GSM-Symbolic数据集中问题文本的整数分布系统性偏向更大数值(K-S统计量=0.12,p<0.001),与原始GSM8K不符。控制这一大数效应后,一半原本显著的案例不再显著。在具有统计显著性能差异的模型中,我们识别出不同的、模型特异性的行为失效模式——包括变量绑定脆弱性、算术限制及双任务干扰,强调对大模型推理能力的泛化断言既存在统计上的过早性,也具有机制误导性。
原文摘要 · Abstract (English)
The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabilities. We argue that this conclusion rests on shaky statistical ground. Re-evaluating 20 open-weight models using bootstrapped Generalised Linear Mixed Models with per-question random effects, we find that only 8 exhibit statistically significant performance changes under the original prompt format. Moreover, we identify a previously unacknowledged factor: the distribution of integers in problem texts of the main GSM-Symbolic dataset is systematically shifted towards larger values relative to the original GSM8K (K-S statistic = 0.12, p < 0.001), contradicting the original authors' claims. Controlling for this large-number effect accounts for significance in half of the remaining cases. Among models with statistically significant performance deltas, we identify distinct, model-specific behavioural failure profiles -- including fragility of variable binding, arithmetic limitations, and dual-task interference -- underscoring that blanket claims about LLM reasoning risk being both statistically premature and mechanistically misleading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。