发现模型谄媚信号在注意力头中线性可分,可针对性干预改善。
Sycophancy Hides Linearly in the Attention Heads
- 通过线性探针定位谄媚信号,发现其在中间层注意力头中最具可分离性。
- 在TruthfulQA上训练的探针可有效迁移至其他事实问答数据集。
- 谄媚行为与求真机制相关但不同,影响头关注用户怀疑表达。
我们发现,正确到错误的谄媚信号在多头注意力激活中具有最强的线性可分性。基于线性表征假设,我们在残差流、多层感知机(MLP)和注意力层上训练线性探针,分析这些信号的出现位置。尽管可分性也出现在残差流和MLP中,但使用这些探针进行引导时,效果最显著的是中间层注意力头的一个稀疏子集。以TruthfulQA为基准数据集,我们发现训练出的探针能有效迁移到其他事实问答基准测试。此外,将发现的方向与先前识别的“求真”方向对比,重叠有限,表明事实准确性与顺从抵抗源自相关但不同的机制。注意力模式分析进一步显示,关键注意力头对用户怀疑表达的关注显著增加,导致谄媚倾向。总体而言,这些发现表明可通过简单的、目标明确的线性干预,利用注意力激活的内部几何结构缓解模型谄媚现象。
原文摘要 · Abstract (English)
We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we train linear probes across the residual stream, multilayer perceptron (MLP), and attention layers to analyze where these signals emerge. Although separability appears in the residual stream and MLPs, steering using these probes is most effective in a sparse subset of middle-layer attention heads. Using TruthfulQA as the base dataset, we find that probes trained on it transfer effectively to other factual QA benchmarks. Furthermore, comparing our discovered direction to previously identified "truthful" directions reveals limited overlap, suggesting that factual accuracy, and deference resistance, arise from related but distinct mechanisms. Attention-pattern analysis further indicates that the influential heads attend disproportionately to expressions of user doubt, contributing to sycophantic shifts. Overall, these findings suggest that sycophancy can be mitigated through simple, targeted linear interventions that exploit the internal geometry of attention activations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。