arXiv:2607.08883cs.LG2026-07AAAI

用激活引导攻击破解大模型安全机制,发现安全表征分布广泛。

Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

论文配图:Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal
图 1 · 摘自论文原文
  • 用内部拒绝方向替代输出目标,直接攻击安全表征
  • 全局抑制各层位置的拒绝行为比单点攻击更有效
  • 新方法提速33倍且提升攻击成功率,适合研究模型安全

大型语言模型的行为对齐常掩盖其脆弱的内部安全表征。近期研究指出,拒绝行为由激活空间中的低维方向所调控。本文以对抗后缀攻击为探针,研究此类表征的结构、定位与优化访问方式。提出激活引导GCG,将输出目标替换为直接针对模型内部拒绝方向的损失。在多种目标变体中发现,全局抑制所有层和位置的拒绝行为比仅针对单一层-位置对更有效,表明安全表征分布在前向传播全程而非局限于单一位置。进一步引入Soft-GCG,通过Gumbel-Softmax实现离散后缀优化的连续松弛,相比标准GCG提速33倍且提升攻击成功率。在不同模型规模下评估显示,小模型仍易受攻击,而大模型在计算受限条件下对激活与后缀攻击均表现出更强抵抗力,符合大模型经更优训练后更难被越狱的预期。结果揭示了当前模型中安全机制的编码与攻破方式,为设计更鲁棒、表征感知的对齐策略提供具体指导。

原文摘要 · Abstract (English)

Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions about how such representations are structured, localized, and accessed by optimization. We study adversarial suffix attacks as a probe of representational alignment. We introduce Activation-Guided GCG, which replaces output-based objectives with losses that directly target a model's internal refusal direction. Across several objective variants, we find that suppressing refusal globally across all layers and positions is more effective than targeting a single layer-position pair. This suggests that safety representations are distributed across the forward pass rather than causally localized to a single site. We further introduce Soft-GCG, a continuous relaxation of discrete suffix optimization using Gumbel-Softmax. Soft-GCG achieves a 33 $\times$ speedup over standard GCG while improving attack success rates. Evaluating across model scales, we find that smaller models remain vulnerable while larger models resist both activation- and suffix-based attacks at our compute-constrained settings, consistent with larger and better safety trained models being harder to jailbreak. Together, our results clarify how safety mechanisms are encoded and can be broken in contemporary models. These insights provide concrete guidance for designing more robust and representation-aware alignment strategies.

模型安全对抗攻击表征分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。