大模型能自我察觉概念注入,且可被引导增强此能力。
Latent Introspection: Models Can Detect Prior Concept Injections
- 通过残差流分析发现模型能检测概念注入
- 引导后检测灵敏度从0.3%升至39.9%
- 适合关注模型安全与内在机制的研究者
我们发现Qwen 32B模型具备潜在的自我觉察能力,能够识别其上下文中的概念注入并定位具体注入内容。尽管模型在生成输出中否认注入,但对残差流的逻辑透镜分析揭示了清晰的检测信号,这些信号在最终层被削弱。当通过提示告知模型关于人工智能自我觉察机制时,检测敏感性显著提升(0.3% → 39.9%),仅带来0.6%的误报率增加。九个注入与恢复概念间的互信息从0.61比特增至1.05比特,排除了噪声解释的可能性。结果表明,模型具有令人意外的内省与导向感知能力,这对潜在推理和安全性具有重要影响。
原文摘要 · Abstract (English)
We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% -> 39.9%) with only a 0.6% increase in false positives. Also, mutual information between nine injected and recovered concepts rises from 0.61 bits to 1.05 bits, ruling out generic noise explanations. Our results demonstrate models can have a surprising capacity for introspection and steering awareness that is easy to overlook, with consequences for latent reasoning and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。