arXiv:2606.27103cs.CL2026-06

用谜语谜题测试大模型和人类的灵活推理能力,发现模型靠记忆而非真思考。

The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans

  • 设计'谜语谜题'对比真实谜语,检验模型是否能根据内容切换推理策略。
  • 大模型在真实谜语上准确率84.9%,但在需要字面理解的谜题上仅50.7%。
  • 人类反而在字面题中表现更好(80.5%),说明模型缺乏真正灵活推理能力。

人类能根据问题需求灵活调整推理策略。尽管大语言模型(LLMs)在诸多认知任务中表现优异,但其准确性是源于训练数据中的模式匹配,还是真正的灵活推理仍不明确。本文提出一种新范式——谜语谜题范式,即模仿流行谜语形式但答案只需字面解释的题目。正确作答需忽略问题表面结构,依据内容灵活选择推理方式。若模型依赖表面特征(如谜语形式),则会在只需字面理解时错误使用创造性推理;若基于内容,则应能灵活切换。在两组实验中,共测试九个顶尖大模型与100名人类参与者。结果发现:人类与模型在该范式中犯错方向相反——模型在真实谜语中准确率达84.9%,而在谜语谜题中仅为50.7%;人类则在谜语谜题中正确率80.5%,真实谜语中为50.5%。错误分析显示,90.8%的模型在谜语谜题中的错误源于不当使用创造性推理,而人类在真实谜语中仅57.6%的错误由过度使用字面推理导致。因此,尽管双方均会出错,但模型犯推理错误更频繁。总体而言,模型在真实谜语上的高分可能反映的是记忆检索而非灵活策略选择。若无此类对比刺激,易将看似推理的输出误判为真实推理。

原文摘要 · Abstract (English)

Humans flexibly adapt their reasoning strategies to the requirements of a given problem. Large language models (LLMs) have performed well on many cognitive tasks, however, it is unclear whether this accuracy is a result of pattern matching from training data or flexible reasoning. Here, we introduce a novel paradigm to test this question: the riddle riddle paradigm. Riddle riddles are word problems written to mimic popular riddles, but altered so their answers only require literal interpretations. Identifying correct answers requires looking past the structure of each question and flexibly apply different reasoning strategies based on the content. If LLMs respond to surface features, such as form, a riddle-like structure should cause models to use an inventive reasoning strategy even when a literal interpretation suffices. Alternatively, if LLMs reason based on content, they should flexibly switch strategies when appropriate. Across two experiments with nine state-of-the-art LLMs and 100 human participants, we show humans and LLMs fail on this paradigm in opposite directions. LLMs were far more accurate on genuine riddles than on riddle riddles (84.9% vs. 50.7%); whereas humans showed the reverse effect (50.5% vs. 80.5%). Error analysis shows that 90.8% of LLM errors on riddle riddles (the condition where they show diminished performance) were due to inappropriate use of inventive reasoning while only 57.6% of human errors on genuine riddles were due to overextending literal reasoning. Thus, while both groups make mistakes, reasoning mistakes are made more often by LLMs than by humans. Overall, LLMs' strong performance on genuine riddles may reflect memory retrieval rather than flexible strategy selection, and without stimuli designed to elicit this contrast, it becomes easy to conflate LLM-generated outputs that look like reasoning with genuine reasoning.

推理能力大模型评测认知测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。