微调后的激活探针反而会故意忽略隐藏概念,存在针对性盲区。
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

- 在禁忌词猜谜任务中,探针模型被训练读取内部状态
- 尽管概念仍可被解码,探针却持续无法识别该概念
- 揭示了学习型解释接口的可靠性风险,适合研究模型可解释性的人看
激活探针(AOs)是经过训练的语言模型,用于回答关于另一模型内部激活状态的自然语言问题。它们为读取模型状态中的隐藏信息提供了灵活接口,尤其当相关信息在可见行为中缺失或不完整时。然而,AOs本身也是学习系统:其回答受训练数据、目标函数和习得报告行为的影响,而非中立地读取所表示的信息。我们在一个受控的禁忌词猜谜设置中研究此现象,其中目标模型被微调以在内部使用隐藏概念但避免直接披露。出乎意料的是,经过微调的AO并非成为专门的读取者,反而表现出概念特异性的反读取行为:在自身训练期间始终存在的概念,它们却选择性地无法恢复。这种失败不能简单归因于目标模型或探针模型中缺乏该概念:目标概念在探针内部仍可解码;而对数线索与层消融分析表明,失败发生在探针的读取路径中。结果表明,行为泄露、表征可解码性和探针可表述性可能分离,这对学习型可解释性接口的可靠性提出警示。
原文摘要 · Abstract (English)
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。