arXiv:2604.01988cs.AI2026-04

测试大模型能否像人一样识别数字规律并合理使用速算技巧。

SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation

  • 设计4800个题目,分8类速算策略和4种数字规模,评估模型用、判、创三能力。
  • 显式提示下模型速算准确率最高提升15%,但默认推理中应用率不足40%。
  • 模型会误用技巧,无法自创符合规则的题目,缺乏对结构的理解。

大型语言模型在存在高效数值捷径时仍倾向于逐步计算,这引发一个基本问题:它们是否具备类似人类的数感,即识别数值结构、适时使用捷径并避免错误使用的能力?我们提出SenseMath,一个可控基准,用于评估大模型在结构敏感型数值推理中的表现。该基准包含4800个样本,覆盖八类捷径和四种数字规模,并设有强捷径、弱捷径与对照组三种变体。支持三种递增认知需求的评估场景:捷径使用(能否在适用问题中应用捷径)、适用性判断(能否识别捷径是否合适或具有误导性)、问题生成(能否从零生成正确支持特定捷径的新题目)。我们在五款大模型(从GPT-4o-mini到Llama-3.1-8B)上进行评估,结果显示:在显式提示下,模型能快速采用捷径策略,在适用题目上准确率最高提升15%;但在标准思维链提示下,自发使用捷径的比例不足40%,即便其已具备相关能力。此外,这种能力仅限于“使用”层面,模型会过度泛化捷径至不适用场景,且无法从头生成有效的捷径相关题目。整体表明当前大模型具备捷径操作熟练度,但缺乏支撑人类数感的结构性理解。

原文摘要 · Abstract (English)

Large language models often default to step-by-step computation even when efficient numerical shortcuts are available. This raises a basic question: do they exhibit number sense in a human-like behavioral sense, i.e., the ability to recognize numerical structure, apply shortcuts when appropriate, and avoid them when they are not? We introduce SenseMath, a controlled benchmark for evaluating structure-sensitive numerical reasoning in LLMs. SenseMath contains 4,800 items spanning eight shortcut categories and four digit scales, with matched strong-shortcut, weak-shortcut, and control variants. It supports three evaluation settings of increasing cognitive demand: Shortcut Use (whether models can apply shortcuts on shortcut-amenable problems); Applicability Judgment (whether they can recognize when a shortcut is appropriate or misleading); and Problem Generation (whether they can generate new problem items that correctly admit a given type of shortcut). Our evaluation across five LLMs, ranging from GPT-4o-mini to Llama-3.1-8B, shows a consistent pattern: when explicitly prompted, models readily adopt shortcut strategies and achieve substantial accuracy gains on shortcut-amenable items (up to 15%), yet under standard chain-of-thought prompting they spontaneously employ such strategies in fewer than 40% of cases, even when they demonstrably possess the requisite capability. Moreover, this competence is confined to the Use level; models systematically over-generalise shortcuts to problems where they do not apply, and fail to generate valid shortcut-bearing problems from scratch. Together, these results suggest that current LLMs exhibit procedural shortcut fluency without the structural understanding of when and why shortcuts work that underlies human number sense.

数感捷径推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。