大模型能察觉自身内部状态变化,识别被注入的概念。
Emergent Introspective Awareness in Large Language Models
- 通过注入概念激活值,测试模型对内部状态的感知能力。
- Claude Opus 4及4.1在多数任务中表现最佳,可识别注入概念并区分自身输出与预填充内容。
- 模型可在指令下主动调控内部表示,具备一定自我监控能力。
我们研究大语言模型是否能对其内部状态进行内省。仅通过对话难以判断真实内省与虚构陈述的区别。为此,我们向模型激活中注入已知概念的表示,并测量这些操作对其自报告状态的影响。结果发现,在特定场景下,模型能够察觉注入概念并准确识别;部分模型具备回忆先前内部表征的能力,可将其与原始文本输入区分开。令人惊讶的是,某些模型能利用对先前意图的记忆,区分自身生成内容与人工预填充内容。所有实验中,性能最强的模型——Claude Opus 4和4.1——表现出最高的内省意识;但跨模型趋势复杂,受后训练策略影响显著。最后,我们探索模型是否能显式控制内部表示,发现当被指令或激励‘思考某个概念’时,模型可调节其激活值。总体而言,当前语言模型对自身内部状态具有一定的功能性内省意识。我们强调,这种能力在现有模型中仍极不可靠且高度依赖上下文;但随着模型能力提升,该能力可能进一步发展。
原文摘要 · Abstract (English)
We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address this challenge by injecting representations of known concepts into a model's activations, and measuring the influence of these manipulations on the model's self-reported states. We find that models can, in certain scenarios, notice the presence of injected concepts and accurately identify them. Models demonstrate some ability to recall prior internal representations and distinguish them from raw text inputs. Strikingly, we find that some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills. In all these experiments, Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness; however, trends across models are complex and sensitive to post-training strategies. Finally, we explore whether models can explicitly control their internal representations, finding that models can modulate their activations when instructed or incentivized to "think about" a concept. Overall, our results indicate that current language models possess some functional introspective awareness of their own internal states. We stress that in today's models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。