重新检验大模型的推理能力,发现其失败多因测试设计问题而非本质缺陷。
Rethinking the Illusion of Thinking
- 通过逐步提示和代理协作对话改进实验方法
- 8个盘子时大模型仍会出错,但100+对代理的渡河问题可轻松解决
- 研究揭示了大模型在复杂推理中的真实局限,适合关注模型机制的研究者
今年早些时候,苹果发布《思维的幻觉》引发争议,批评者认为大推理模型缺乏真正推理能力,仅是随机鹦鹉。支持者则指出实验设计存在缺陷。本文复现并优化了其中最具争议的两个基准测试:汉诺塔与渡河问题。通过引入分步提示和代理协作对话,我们发现汉诺塔的失败不仅源于输出限制,也反映认知局限——当盘数增至约8个时模型便易出错。而最初被视作灾难性失败的渡河问题,实因测试了无解配置所致。严格限定为可解问题后,模型可轻松解决超过100对代理的复杂实例。研究结论表明,当前大推理模型本质上是离散状态空间中经强化学习调优的搜索器,其符号化长程推理进展需通过精细消融分析推进。
原文摘要 · Abstract (English)
Earlier this year, Apple ignited controversy by publishing "The Illusion of Thinking," prompting heated debate within the AI community. Critics seized upon the findings as conclusive evidence that Large Reasoning Models (LRMs) lack genuine reasoning capabilities, branding them as mere stochastic parrots. Meanwhile, defenders-spearheaded by Lawsen et al. (2025)-fired back, condemning the experimental setup as flawed and the conclusions overstated. We clarify this debate by replicating and refining two of the original study's most contentious benchmarks: Towers of Hanoi and River Crossing. By introducing incremental stepwise prompting and agentic collaborative dialogue, we show that previously reported failures solving the Towers of Hanoi were not purely result of output constraints, but also partly a result of cognition limitations: LRMs still stumble when complexity rises moderately (around 8 disks). Moreover, the River Crossing results initially heralded as catastrophic failures turn out to hinge upon testing unsolvable configurations. Once we limit tests strictly to solvable problems-LRMs effortlessly solve large instances involving over 100 agent pairs. Our findings ultimately defy simplistic narratives: today's LRMs are stochastic, RL-tuned searchers in a discrete state space we barely understand. Real progress in symbolic, long-horizon reasoning demands mapping that terrain through fine-grained ablations like those introduced here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。