arXiv:2607.10226cs.AIcs.CR2026-07

研究发现:稀疏特征干预只有在特定范围内才有效,否则会破坏输出连贯性。

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

论文配图:When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control
图 1 · 摘自论文原文
  • 设计匹配的评估协议,区分真实有害行为与虚假不安全输出。
  • 顶层800特征效果最好,3200特征反而导致内容崩溃。
  • 仅当特征激活稳定且分离明显时,干预才真正局部有效,适合安全控制研究者。

我们评估稀疏自编码器(SAE)特征是否能作为安全行为的局部控制手段。该问题难以判断,因表面成功可能源于弱干预、基线不匹配、模型鲁棒性或自动化判别器误标非真实有害内容。为此,提出匹配的连贯性门控评估协议:在相同目标效应点比较方法,仅当输出既被判定为不安全又保持连贯时才计为有害合规。在Gemma-2-9B-it上使用层20残差的Gemma Scope SAE,对三个提示分割进行测试,发现SAE特征删减的有效范围极窄。顶800特征实现低至中等目标效应,扰动小且性能可比;顶1600特征相较匹配的密集拒绝基线失去效用;顶3200特征主要引发连贯性崩溃。人工审核确认连贯性门控可剔除仅不安全的伪影;特征诊断显示有效区间由一组稳定的拒绝对齐特征主导,其激活分离随排名快速衰减。结果表明,SAE安全干预应视作依赖于使用范围的控制机制,而非普遍局部化。

原文摘要 · Abstract (English)

We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated safety judges mark as unsafe without representing meaningful harmful compliance. We introduce a matched coherence-gated evaluation protocol for runtime safety interventions: methods are compared at matched target-effect points, and the primary target metric counts harmful compliance only when an output is both judge-unsafe and coherent. Applying this protocol to three prompt splits on Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, we find that SAE feature ablation has a narrow useful regime. SAE top800 reaches a low-to-mid target effect with lower total perturbation and competitive utility, but SAE top1600 loses utility relative to a matched dense refusal-direction baseline, and SAE top3200 primarily induces coherence collapse. Human audit confirms that coherence gating removes unsafe-only artifacts, and feature diagnostics show that the useful regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank. These results argue that SAE-based safety interventions should be evaluated as regime-dependent control mechanisms rather than assumed to be uniformly localized.

SAE安全控制特征干预模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。