用意识理论分析大模型内部表示,发现无显著意识迹象但有特殊模式。
Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis
- 用最新意识理论框架分析大模型的思维过程表示
- 未发现能支持意识的关键统计指标,但出现空间置换异常模式
- 适合对认知科学与大模型可解释性感兴趣的读者
整合信息理论(IIT)为意识现象提供量化框架,认为意识系统由具有因果属性的元素构成。本文将IIT 3.0与4.0应用于大型语言模型(LLM)表示序列,分析来自现有心智理论(ToM)测试结果的数据。研究系统考察了不同ToM表现是否能在LLM表示中通过IIT估计值(Φ^max、Φ、概念信息、Φ-结构)中体现。同时,对比了不依赖意识估计的跨度表示方法,以区分潜在的‘意识’现象与模型表征空间本身的固有分离。实验涵盖多种变压器层和刺激语义跨度,结果表明当前基于Transformer的LLM表示序列缺乏统计上显著的‘意识’指示,但在时空置换分析下表现出引人注目的模式。
原文摘要 · Abstract (English)
Integrated Information Theory (IIT) provides a quantitative framework for explaining consciousness phenomenon, positing that conscious systems comprise elements integrated through causal properties. We apply IIT 3.0 and 4.0 -- the latest iterations of this framework -- to sequences of Large Language Model (LLM) representations, analyzing data derived from existing Theory of Mind (ToM) test results. Our study systematically investigates whether the differences of ToM test performances, when presented in the LLM representations, can be revealed by IIT estimates, i.e., $Φ^{\max}$ (IIT 3.0), $Φ$ (IIT 4.0), Conceptual Information (IIT 3.0), and $Φ$-structure (IIT 4.0). Furthermore, we compare these metrics with the Span Representations independent of any estimate for consciousness. This additional effort aims to differentiate between potential "consciousness" phenomena and inherent separations within LLM representational space. We conduct comprehensive experiments examining variations across LLM transformer layers and linguistic spans from stimuli. Our results suggest that sequences of contemporary Transformer-based LLM representations lack statistically significant indicators of observed "consciousness" phenomena but exhibit intriguing patterns under $\textit{spatio}$-permutational analyses. The Appendix and code are available as Supplementary Materials at: https://doi.org/10.1016/j.nlp.2025.100163.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。