揭秘大模型拒绝行为操控的内在机制,发现注意力核心路径。
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
- 通过多标记激活修补法,定位到操纵拒绝行为的关键电路。
- 仅8.83%性能下降,证明注意力分数冻结不影响效果。
- 可压缩96%向量维度,且不同方法共享关键特征维度。
对大型语言模型(LLM)应用操控向量是一种高效且有效的对齐技术,但其内部工作机制尚不明确——具体而言,操控向量如何影响模型内部结构并导致输出差异仍不清楚。为探究操控向量有效性的因果机制,本文以拒绝行为为案例进行系统研究。提出一种多标记激活修补框架,发现不同操控方法在相同层中依赖功能可互换的电路。这些电路显示,操控向量主要通过OV电路作用于注意力机制,而几乎不涉及QK电路。在操控过程中冻结所有注意力得分,仅导致三个模型家族平均性能下降8.83%。对被操控的OV电路进行数学分解,揭示出语义可解释的概念,即使原始操控向量本身不可解释。基于激活修补结果,发现操控向量可稀疏化达85%-96%仍保持大部分性能,且不同操控方法在重要维度上达成一致。
原文摘要 · Abstract (English)
Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works--specifically, what internal mechanisms steering vectors affect and how this results in different model outputs. To investigate the causal mechanisms underlying the effectiveness of steering vectors, we conduct a comprehensive case study on refusal. We propose a multi-token activation patching framework and discover that different steering methodologies leverage functionally interchangeable circuits when applied at the same layer. These circuits reveal that steering vectors primarily interact with the attention mechanism through the OV circuit while largely ignoring the QK circuit. Freezing all attention scores during steering drops performance by only 8.83% across three model families. A mathematical decomposition of the steered OV circuit further reveals semantically interpretable concepts, even in cases where the steering vector itself does not. Leveraging the activation patching results, we show that steering vectors can be sparsified by up to 85-96% while retaining most performance, and that different steering methodologies agree on a subset of important dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。