用机制可解释性揭示大模型理解的三层能力,打破‘仅模仿’的刻板印象。
Mechanistic Indicators of Understanding in Large Language Models
- 分三层次定义理解:概念、世界状态、原理级,对应不同层级的内部表征
- 发现模型在隐空间中形成特征方向,能统一不同表现形式的同一实体
- 适合研究大模型认知机制、哲学与人工智能交叉的学者参考
大型语言模型常被视作仅模仿语言模式而无真正理解。我们指出,机制可解释性(MI)的新进展表明这一观点已难成立——但前提是将这些发现整合进理解的理论框架。本文提出一个分层框架,区分三种理解层级:当模型在隐空间中形成“特征”方向,建立同一实体或属性的不同表现之间的联系时,出现概念理解;当模型学习特征间的偶然事实关联并动态追踪世界变化时,出现世界状态理解;当模型不再依赖记忆事实,而是发现连接事实的紧凑“电路”时,出现原则理解。在此框架下,MI揭示了可支撑类理解统一性的内部组织结构。然而,这些结构也与人类认知存在差异,表现为并行利用异质机制。结合哲学理论与机制证据,我们超越了‘是否理解’的二元争论,推动一种基于机制比较的进化认识论,探索人工智能理解与人类理解的共性与分歧。
原文摘要 · Abstract (English)
Large language models (LLMs) are often portrayed as merely imitating linguistic patterns without genuine understanding. We argue that recent findings in mechanistic interpretability (MI), the emerging field probing the inner workings of LLMs, render this picture increasingly untenable--but only once those findings are integrated within a theoretical account of understanding. We propose a tiered framework for thinking about understanding in LLMs and use it to synthesize the most relevant findings to date. The framework distinguishes three hierarchical varieties of understanding, each tied to a corresponding level of computational organization: conceptual understanding emerges when a model forms "features" as directions in latent space, learning connections between diverse manifestations of a single entity or property; state-of-the-world understanding emerges when a model learns contingent factual connections between features and dynamically tracks changes in the world; principled understanding emerges when a model ceases to rely on memorized facts and discovers a compact "circuit" connecting these facts. Across these tiers, MI uncovers internal organizations that can underwrite understanding-like unification. However, these also diverge from human cognition in their parallel exploitation of heterogeneous mechanisms. Fusing philosophical theory with mechanistic evidence thus allows us to transcend binary debates over whether AI understands, paving the way for a comparative, mechanistically grounded epistemology that explores how AI understanding aligns with--and diverges from--our own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。