通过潜空间攻击绕过语言模型的拒绝机制,提升恶意请求成功率
Latent-space Attacks for Refusal Evasion in Language Models

- 将拒绝规避建模为对线性探测器的潜空间逃避攻击
- 新方法在15个模型上实现最高攻击成功率,超越现有基线
- 适合研究安全漏洞、对抗攻击或模型可解释性的研究人员
安全对齐的语言模型会拒绝有害请求,但可通过操纵其内部表示抑制拒绝行为。现有方法通过从模型激活中移除拒绝方向来实现,旨在消除残差流中的拒绝信号。尽管实验成功,但缺乏对所诱导潜空间变换的理论解释。本文将拒绝抑制重新定义为针对分离被拒与被答提示的线性探测器的潜空间逃避攻击。先前工作的均值差异方向自然构成此类探测器,其消融即为投影至决策边界,相当于最小置信度逃避攻击。该视角不仅解释了先前方法的成功,还揭示关键局限:逃避止于决策边界,因此需将表示进一步推入模型回答区域。为此,我们提出受控潜空间逃避攻击,通过优化置信度将表示投影过边界。在15个指令微调、多模态和推理模型上,该方法达到最优攻击成功率,优于现有拒绝消融基线及专用越狱攻击。
原文摘要 · Abstract (English)
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by ablating a refusal direction from model activations, aiming to remove refusal from the model's residual stream. Despite their empirical success, these methods lack a principled account of the latent-space transformation they induce and why it suppresses refusal. In this work, we recast refusal suppression as a latent-space evasion attack against linear probes trained to separate refused from answered prompts. Under this view, prior work's difference-in-means direction naturally defines such a probe, and its ablation is exactly a projection onto its decision boundary, i.e., a minimum-confidence evasion attack. This perspective not only explains the empirical success of prior work but also admits a key limitation: evasion stops at the decision boundary, motivating the need to push representations further into the compliant region, i.e., where the model answers. We leverage this by proposing a Controlled Latent-space Evasion attack that projects representations past the boundary with an optimized confidence. We achieve state-of-the-art attack success rate across 15 instruction-tuned, multimodal, and reasoning models, outperforming existing refusal-ablation baselines and specialized jailbreak attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。