arXiv:2604.07593cs.AI2026-04

长题干和长解答会增加大模型出错概率,尤其在高难度数学题中。

Too long; didn't solve

  • 分析题干与解题过程长度对模型表现的影响
  • 长度越长,模型错误率越高,且在困难题目中更明显
  • 适合关注模型推理鲁棒性与评测设计的研究者

数学基准测试广泛用于评估大语言模型的推理能力,但其结构特性如何影响模型行为仍不明确。本文研究了题干长度与解答长度两个结构变量,分析它们与新构建的专家编写数学难题对抗数据集上模型表现的关系。结果发现,题干和解答长度均与模型失败率呈正相关。此外,进行跨模型分歧的探索性分析,在难度归一化后,两者与模型实际差异仍保持微弱负相关,题干长度关联稍强。总体而言,结构长度与该数据集中的实证难度密切相关。

原文摘要 · Abstract (English)

Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language models, yet little is known about how their structural properties influence model behaviour. In this work, we investigate two structural length variables, prompt length and solution length, and analyse how they relate to model performance on a newly constructed adversarial dataset of expert-authored mathematics problems. We find that both prompt and solution lengths correlate positively with increased model failure across models. We also include a secondary, exploratory analysis of cross-model disagreement. Under a difficulty-adjusted normalised analysis, both variables retain weak negative associations with realised model separation, slightly stronger for prompt length. Overall, our main robust finding is that structural length is linked to empirical difficulty in this dataset.

数学推理模型评测结构长度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。