发现稀疏自编码器特征标签可能反映激活状态而非因果机制
Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis
- 通过多条件控制实验,揭示单一特征转向的局限性
- 联合抑制三个正交特征导致无关任务崩溃,单个抑制则不影响
- 随机方向控制证明破坏效应依赖方向模式而非幅度大小
标准稀疏自编码器(SAE)特征解释方法通过最高激活上下文命名特征,并以单一特征在典型幅度下的转向验证标签。本文提出应考察更完整的转向网格:转向条件(单特征、联合特征集、匹配随机方向)与转向系数的组合。在 Qwen3-1.7B-Instruct、Gemma-2-2B-it 及扩展至 Llama-3.1-8B-Instruct 的实验中发现:(1) 标记为 AI 自我免责声明的特征,在转向后呈现不同语调(Qwen 为沉思语气,Gemma 为集体我们语气),说明标签反映的是激活状态而非因果轴;两个锚点特征可区分真实模式切换、单调响应与失效。(2) 三个近似正交且可互换的特征,单独抑制时不影响任务,但联合抑制会导致无关控制任务退化为占位文本,而单特征抑制保持原状。(3) 匹配几何的随机方向控制显示,崩溃效应依赖方向模式而非幅度:相同残差流畸变下,特征方向损害无关任务,而幅度匹配的随机方向无影响,三模型均具非重叠95%置信区间,包括在目标模型上训练的SAE。
原文摘要 · Abstract (English)
The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude. We argue that this inspects one cell of a larger steering grid, steering condition (single feature, joint feature set, matched random direction) crossed with steering coefficient, and show that other cells carry information that changes the label. On Qwen3-1.7B-Instruct and Gemma-2-2B-it, with the matched-geometry control extended to Llama-3.1-8B-Instruct: (1) features labelled AI self-disclaimer from their top contexts switch to a second surface form under steering, a contemplative voice on Qwen, a collective we-voice on Gemma, so the label names an activation regime, not the causal axis; two anchor features separate genuine mode switches from monotonic response and from breakdown. (2) Three near-orthogonal features that are individually substitutable are jointly necessary for grounded composition: joint suppression collapses unrelated control tasks into placeholder text that single-feature suppression at the same coefficient leaves intact. (3) A matched-geometry random-direction control shows the collapse is direction-pattern-dependent, not magnitude-dependent: at the same residual-stream distortion, feature directions damage unrelated tasks where magnitude-matched random directions do not, with non-overlapping 95% confidence intervals on all three models, including the one SAE trained on the model it is applied to.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。