SAE干预看似有效,实则行为可恢复,安全防护存在隐藏漏洞。
SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

- 通过优化残差扰动,验证干预后行为仍可恢复
- 在拒绝生成任务中恢复率达95.8%,特征漂移仅0.131
- 适合关注AI安全与可控性研究的读者
稀疏自编码器(SAEs)将残差流激活分解为可解释特征。近期潜在空间防御依赖这些特征,假设识别出的“危险”特征可作为监控与干预的操作点。然而,我们发现该方法存在可恢复的失败模式:钳制特定有害特征仅阻断可见路径,并未真正消除行为。我们将其建模为后干预恢复问题,即在干预后残差状态基础上,优化残差扰动以恢复原始行为,同时保持目标SAE特征值不变。即使在干预全程活跃的强威胁模型下,恢复仍可实现。为排除干预被简单撤销,采用编码器正交更新(单层)和跨层特征映射雅可比矩阵。在TPP、遗忘、IOI及拒绝引导实验中,均发现行为可恢复。尤其在安全关键的拒绝引导任务中,对有效样本的恢复率达95.8%,而受保护特征相对漂移仅为0.131,显著低于基于后缀的基线。恢复路径归因分析进一步将根源定位至SAE重构残差,即未被SAE解释的部分。结果揭示特征级控制与行为完整性之间存在差距:虽可支持因果干预,但控制特征并不保证控制行为。
原文摘要 · Abstract (English)
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that identified "unsafe" SAE features serve as actionable handles for monitoring and intervention. In this paradigm, clamping a specific harmful feature is expected to reliably prevent model misbehavior. However, we show that this success may hide a recoverable failure mode: the clamp may block one visible route to a behavior without eliminating the behavior itself. We formulate this vulnerability as post-intervention recovery, a constrained residual-space optimization problem. Starting from the post-intervention residual state, we optimize residual perturbations to recover the pre-intervention behavior while preserving the post-intervention values of the targeted SAE features. Even under a strong threat model where the intervention remains active throughout optimization and generation, recovery remains possible. To rule out that recovery simply undoes the intervention, we use encoder-orthogonal updates for single-layer interventions and the corresponding feature-map Jacobian in the cross-layer setting. Across TPP, unlearning, IOI, and refusal steering experiments, this stress test reveals recoverable behavior despite successful feature-level intervention. Especially in the safety-critical refusal-steering setting, we achieve a 95.8% recovery rate on valid samples while keeping defended-feature relative drift to 0.131, substantially below suffix-based baselines. A recovery-path attribution analysis further localizes this recovery to the SAE reconstruction residual, the component left unexplained by the SAE. These results expose a gap between feature-level control and behavioral completeness: SAE features can support causal intervention, but controlling them does not guarantee control over the underlying behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。