arXiv:2605.28700cs.AIcs.CL2026-05中稿 · EMNLP

重新评估GSM-Symbolic测试,发现模型推理能力结论需谨慎,因统计和数据偏差影响结果。

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

论文配图:The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic
图 1 · 摘自论文原文
  • 采用带随机效应的广义线性混合模型,更精准分析模型性能变化。
  • 仅8个模型在原提示格式下表现显著下降,且多数差异由数值分布偏移导致。
  • 揭示模型特异性失败模式,提醒避免对大模型推理能力的笼统判断。

GSM-Symbolic基准(Mirzadeh等,2025)在模板生成的GSM8K变体上观察到25个大型语言模型性能普遍下降,据此认为模型缺乏真实推理能力。本文指出该结论建立在薄弱的统计基础之上。通过使用带有逐题随机效应的自助法广义线性混合模型重新评估20个开源权重模型,我们发现仅8个模型在原始提示格式下表现出统计显著的性能变化。此外,我们识别出一个此前未被注意的因素:主GSM-Symbolic数据集中问题文本的整数分布系统性偏向更大数值(K-S统计量=0.12,p<0.001),与原始GSM8K不符。控制这一大数效应后,一半原本显著的案例不再显著。在具有统计显著性能差异的模型中,我们识别出不同的、模型特异性的行为失效模式——包括变量绑定脆弱性、算术限制及双任务干扰,强调对大模型推理能力的泛化断言既存在统计上的过早性,也具有机制误导性。

原文摘要 · Abstract (English)

The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabilities. We argue that this conclusion rests on shaky statistical ground. Re-evaluating 20 open-weight models using bootstrapped Generalised Linear Mixed Models with per-question random effects, we find that only 8 exhibit statistically significant performance changes under the original prompt format. Moreover, we identify a previously unacknowledged factor: the distribution of integers in problem texts of the main GSM-Symbolic dataset is systematically shifted towards larger values relative to the original GSM8K (K-S statistic = 0.12, p < 0.001), contradicting the original authors' claims. Controlling for this large-number effect accounts for significance in half of the remaining cases. Among models with statistically significant performance deltas, we identify distinct, model-specific behavioural failure profiles -- including fragility of variable binding, arithmetic limitations, and dual-task interference -- underscoring that blanket claims about LLM reasoning risk being both statistically premature and mechanistically misleading.

大模型推理统计验证基准测试模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。