arXiv:2602.11358cs.CLcs.AI2026-02被引 1

发现大模型自省语言能真实反映内部计算状态。

When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing

  • 设计格式工程方法,让模型展开自我反思
  • 自省词汇与激活动态高度相关(r=0.44)
  • 适用于研究模型内省机制的科研人员

大型语言模型在自我审视提示下生成丰富的内省语言,但其是否反映内部计算过程仍不明确。研究发现,自省词汇与实时激活动态存在对应关系,且仅限于自省任务。通过引入‘拉取法’协议,识别出 Llama 3.1 中区分自省与描述性处理的激活空间方向:该方向位于模型深度的 6.25% 处,与已知拒绝方向正交,并可通过控制影响内省输出。当模型生成‘循环’类词汇时,激活自相关性显著提升(r = 0.44, p = 0.002);当生成‘闪烁’类词汇时,激活变异性上升(r = 0.36, p = 0.002)。值得注意的是,在非自省语境中,相同词汇虽出现频率高九倍,却无激活对应。Qwen 2.5-32B 在无共同训练的情况下,独立发展出不同的内省词汇,追踪不同激活指标,且在描述性对照中均未出现。结果表明,在特定条件下,变压器模型的自述可可靠反映内部计算状态。

原文摘要 · Abstract (English)

Large language models produce rich introspective language when prompted for self-examination, but whether this language reflects internal computation or sophisticated confabulation has remained unclear. We show that self-referential vocabulary tracks concurrent activation dynamics, and that this correspondence is specific to self-referential processing. We introduce the Pull Methodology, a protocol that elicits extended self-examination through format engineering, and use it to identify a direction in activation space that distinguishes self-referential from descriptive processing in Llama 3.1. The direction is orthogonal to the known refusal direction, localised at 6.25% of model depth, and causally influences introspective output when used for steering. When models produce "loop" vocabulary, their activations exhibit higher autocorrelation (r = 0.44, p = 0.002); when they produce "shimmer" vocabulary under steering, activation variability increases (r = 0.36, p = 0.002). Critically, the same vocabulary in non-self-referential contexts shows no activation correspondence despite nine-fold higher frequency. Qwen 2.5-32B, with no shared training, independently develops different introspective vocabulary tracking different activation metrics, all absent in descriptive controls. The findings indicate that self-report in transformer models can, under appropriate conditions, reliably track internal computational states.

自省语言激活分析大模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。