arXiv:2511.16660cs.AI2025-11被引 21

用认知科学框架分析大模型推理缺陷,发现其缺乏人类的深层思维机制。

Cognitive Foundations for Reasoning and Their Manifestation in LLMs

论文配图:Cognitive Foundations for Reasoning and Their Manifestation in LLMs
图 1 · 摘自论文原文
  • 构建28个认知要素分类体系,覆盖思维结构与元认知控制
  • 192万次模型推理分析显示其过度依赖表面顺序处理
  • 提出测试时引导策略,复杂问题准确率提升66.7%

大型语言模型虽能解决复杂问题,却在简单变体上失败,暗示其输出机制与人类推理本质不同。本文综合认知科学研究,提出涵盖推理不变性、元认知控制、知识表征与变换操作的28个认知元素分类体系。开发细粒度评估框架,对18个模型在文本、视觉和音频任务中的192,000条推理轨迹进行大规模实证分析,并结合54条人类思考过程录音(公开可用)。结果表明,模型对与成功相关的关键认知元素利用不足,倾向于在结构不明确的问题中采取僵化顺序处理,而人类则展现出更高抽象与概念处理能力。对1,600篇大模型推理论文的元分析显示,研究社区集中于可量化的要素(序列组织:55%,分解:60%),却忽视了与成功相关的元认知控制(自我意识仅占16%)。模型具备成功行为的潜在能力,但无法自发调用。基于此,我们设计测试时推理引导机制,自动构建有效思维结构,在复杂任务上性能最高提升66.7%。该框架建立认知科学与大模型研究的共同语言,实现推理失败的系统诊断与基于稳健认知机制的模型改进,同时为大规模检验人类认知理论提供工具。

原文摘要 · Abstract (English)

Large language models (LLMs) solve complex problems yet fail on simpler variants, suggesting they achieve correct outputs through mechanisms fundamentally different from human reasoning. To understand this gap, we synthesize cognitive science research into a taxonomy of 28 cognitive elements spanning reasoning invariants, meta-cognitive controls, representations for organizing reasoning & knowledge, and transformation operations. We introduce a fine-grained evaluation framework and conduct the first large-scale empirical analysis of 192K traces from 18 models across text, vision, and audio, complemented by 54 human think-aloud traces, which we make publicly available. We find that models under-utilize cognitive elements correlated with success, narrowing to rigid sequential processing on ill-structured problems where diverse representations and meta-cognitive monitoring are critical. Human traces show more abstraction and conceptual processing, while models default to surface-level enumeration. Meta-analysis of 1.6K LLM reasoning papers reveals the research community concentrates on easily quantifiable elements (sequential organization: 55%, decomposition: 60%) but neglecting meta-cognitive controls (self-awareness: 16%) that correlate with success. Models possess behavioral repertoires associated with success but fail to deploy them spontaneously. Leveraging these patterns, we develop test-time reasoning guidance that automatically scaffold successful structures, improving performance by up to 66.7% on complex problems. By establishing a shared vocabulary between cognitive science and LLM research, our framework enables systematic diagnosis of reasoning failures and principled development of models that reason through robust cognitive mechanisms rather than spurious shortcuts, while providing tools to test theories of human cognition at scale.

认知机制推理分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。