模型不是不会思考,而是缺工具执行,换个方式就能突破认知极限。
A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- 用智能体工具替代纯文本生成,打破推理悬崖
- 同一模型在工具支持下解决复杂谜题并应对更高难度变体
- 揭示了推理能力与执行能力的差距,适合关注AI行动力的研究者
Shojaee等人(2025)提出,大推理模型(LRMs)在问题复杂度超过某阈值后性能骤降,称为‘推理悬崖’,并认为这是链式思维(CoT)的固有缩放限制。本文反驳该结论,指出其结果受实验设计缺陷干扰:静态文本评估范式中存在工具使用受限、上下文记忆不足、缺乏认知基线、统计报告不全及输出生成瓶颈等系统性约束。我们重新从‘智能体缺口’视角解读该现象,认为模型并非推理失败,而是因接口过于封闭而无法执行。实验证明,原本宣称无法解决的谜题,在启用智能体工具后被成功攻克,并能处理远超原‘推理悬崖’复杂度的变体。对o4-mini和GPT-4o等工具增强模型的分析显示,存在从简单流程执行到复杂元认知自我修正的智能体推理层级,这对机器智能的定义与评测具有深远意义。所谓‘思考幻觉’本质是具备能力却无行动手段所致。
原文摘要 · Abstract (English)
The recent work by Shojaee et al. (2025), titled The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity, presents a compelling empirical finding, a reasoning cliff, where the performance of Large Reasoning Models (LRMs) collapses beyond a specific complexity threshold, which the authors posit as an intrinsic scaling limitation of Chain-of-Thought (CoT) reasoning. This commentary, while acknowledging the study's methodological rigor, contends that this conclusion is confounded by experimental artifacts. We argue that the observed failure is not evidence of a fundamental cognitive boundary, but rather a predictable outcome of system-level constraints in the static, text-only evaluation paradigm, including tool use restrictions, context window recall issues, the absence of crucial cognitive baselines, inadequate statistical reporting, and output generation limits. We reframe this performance collapse through the lens of an agentic gap, asserting that the models are not failing at reasoning, but at execution within a profoundly restrictive interface. We empirically substantiate this critique by demonstrating a striking reversal. A model, initially declaring a puzzle impossible when confined to text-only generation, now employs agentic tools to not only solve it but also master variations of complexity far beyond the reasoning cliff it previously failed to surmount. Additionally, our empirical analysis of tool-enabled models like o4-mini and GPT-4o reveals a hierarchy of agentic reasoning, from simple procedural execution to complex meta-cognitive self-correction, which has significant implications for how we define and measure machine intelligence. The illusion of thinking attributed to LRMs is less a reasoning deficit and more a consequence of an otherwise capable mind lacking the tools for action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。