arXiv:2607.01006cs.CL2026-07被引 7

解析大模型认知机制,揭示其类人能力与本质差异。

Understanding Large Language Models

  • 基于Transformer与注意力机制,实现海量数据下的通用语言建模。
  • 实证显示大模型具备符号推理、心智理论等类人能力,也暴露记忆依赖缺陷。
  • 通过可解释性分析揭示认知机制,呼吁超越简单类比的理性讨论。

大语言模型(LLMs)是近年来人工智能与自然语言处理领域最重要的进展之一。尽管如此,其工作机制、能力边界及与人类认知的关系仍存在诸多争议。本文通过梳理近期证据,探讨了大模型涌现的能力及其在各处理层中的实现机制。首先简要介绍Transformer架构,强调注意力机制如何使模型在大规模数据上训练,从而成为通用而非专用模型。接着分析大模型表现出的类人认知能力,如符号推理、心智理论和欺骗策略,多项研究证实其能解决曾被认为需人类认知的任务;同时也有研究揭示其失败案例,凸显人与大模型认知的本质差异。此外,本文回顾了从神经元激活分析到电路追踪的可解释性方法。最后,针对大模型是否真正理解展开讨论:反对拟人化的观点认为其行为源于对训练数据的模式记忆,而非真实认知。我们指出这一观点受制于对优化过程与认知容量的误解,主张以更细致的视角探讨大模型认知,既承认人机差异,也拒绝以简单还原论否定人工智能认知的可能性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) represent one of the most significant advances in AI and natural language processing in recent years. Still, many pressing questions about their mechanisms, capabilities, and relationship to human cognition remain highly debated. This chapter aims to outline our current understanding of LLMs by discussing recent evidence on emerging capabilities and their mechanistic implementation within processing layers. We begin with a concise overview of the Transformer architecture, emphasizing how the attention mechanism enables training on massive datasets, allowing LLMs to function as generalist rather than specialized models. Next, we examine emergent LLM capabilities that appear to resemble aspects of human cognition, including symbolic reasoning, theory of mind, and deception strategies. Several studies provide evidence that LLMs can solve tasks previously thought to require human-like cognition. Other studies reveal insightful failure cases that shed light on the differences between human and LLM cognition. Alongside these findings, we review explainable AI approaches ranging from neuron activation analysis to circuit tracing. In the final section, we address current debates concerning what LLMs genuinely understand versus what they merely appear to understand. Prominent arguments against AI anthropomorphism point to the simplicity of LLM training objectives, claiming that LLM behavior is better explained by pattern memorization of training data than by genuine cognition. We argue that this standpoint is guided by misconceptions about optimization processes and cognitive capacity, and advocate for a more nuanced discussion of LLM cognition that neither dismisses the differences between humans and LLMs nor precludes the possibility of AI cognition through overly simplistic reductionist arguments.

大模型认知机制可解释性类人能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。