arXiv:2510.11812cs.CLcs.AI2025-10

测试大模型在变形逻辑题上的真实推理能力,发现它们常凭记忆答题。

PHANTOM RECALL: When Familiar Puzzles Fool Smart Models

  • 设计25个经典谜题及149种细节变化的变体,保持逻辑结构不变
  • 多数模型在题目微调后准确率暴跌,仍自信输出错误答案
  • 提出检测工具和提示优化框架,帮助模型真正重新思考问题

大型语言模型(如GPT、Gemini、Claude)看似擅长解答经典逻辑谜题,但其背后是否具备真正的推理能力?近期证据表明,这些模型常依赖记忆模板而非从原理出发推理。当谜题稍作修改时,其表现急剧下降,暴露显著脆弱性。为此,我们系统研究:大模型是否已解决此问题?程度如何?对其他谜题的扰动又如何?是否存在通用提示重写方式提升表现?为此,我们提出PHANTOM RECALL基准,包含25个知名逻辑谜题及149个精心设计的扰动版本,保持推理结构不变但改变表面细节与解法。评估11个主流大模型,发现普遍存在‘幻觉回忆’现象——模型自信复现记忆中的解法或无关推理解释,不再适配新场景。为深入分析并缓解该问题,我们贡献三个工具:(i) 自动化逻辑等价性判断器以检测推理偏差,(ii) 细粒度推理错误分类体系,(iii) 基于分类指导的提示缓解框架。尽管在原始谜题上接近完美准确率,模型在扰动谜题上远低于人类表现,表现出幻觉回忆与过度推演双重缺陷。研究揭示关键局限:当上下文线索变化时,模型往往无法重新推理——凸显语言流畅性与逻辑理解之间的鸿沟。

原文摘要 · Abstract (English)

Large language models (LLMs) such as GPT, Gemini, and Claude often appear adept at solving classic logic puzzles--but how much genuine reasoning underlies their answers? Recent evidence suggests that these models frequently rely on memorized templates rather than reasoning from first principles. When puzzles are slightly modified, their performance collapses, revealing a striking fragility. In particular, we asked: Have LLMs addressed these issues? To what extent? How about perturbations to other puzzles? Is there a general way of reformulating the prompt so that the models do better? To examine these things systematically, we introduce PHANTOM RECALL, a benchmark comprising 25 well-known logic puzzles and 149 carefully designed perturbations that preserve reasoning structure but alter superficial details and solutions. We evaluate eleven leading LLMs and identify a recurring failure mode--phantom recall--where models confidently reproduce memorized solutions or spurious rationales that no longer fit the altered scenario. To probe and mitigate this issue, we contribute three tools: (i) an automated logical-equivalence judge to detect reasoning mismatches, (ii) a taxonomy of fine-grained reasoning error categories, and (iii) a prompting-based mitigation framework guided by these categories. Despite near-perfect accuracy on unmodified puzzles, models significantly underperform humans on perturbed ones, exhibiting both phantom recall and over-elaboration. Our findings reveal a crucial limitation: LLMs often fail to re-reason when contextual cues shift--highlighting the gap between linguistic fluency and logical understanding.

大模型逻辑推理幻觉提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。