发现拒绝行为的动态轨迹,用轻量方法提升越狱攻击检测能力
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection

- 通过因果追踪定位拒绝行为的动态激活路径
- 在多个大模型上提升越狱攻击检测准确率,即使攻击压制了最终拒绝信号
- 适合安全研究者与模型防御开发者使用
现有表示工程通常用静态方向表征拒绝行为,但忽略了其在层-词元位置上的动态构建过程。我们通过因果追踪识别出一种'拒绝轨迹':一种稀疏的上游激活模式,即使在GCG等攻击压制终端拒绝信号时仍能持续存在。基于此,提出SALO(稀疏激活定位算子),一种轻量级白盒检测器,直接作用于选定层窗口的原始隐藏状态。在Qwen、Llama和Mistral模型上,SALO在固定XSTest校准阈值下,对多种攻击家族均实现检测性能提升。进一步分析表明,静态RepE类基线存在局限性,且输入编码边界情况影响显著,揭示了拒绝轨迹监控的潜力与边界。
原文摘要 · Abstract (English)
Representation Engineering analyses often characterize refusal using static directions extracted from terminal or pooled representations. We ask whether this view misses how refusal is constructed across layer-token positions. Using causal tracing, we identify a \textit{Refusal Trajectory}: a sparse upstream activation pattern that often persists even when attacks such as GCG suppress terminal refusal signals. Based on this observation, we propose SALO (Sparse Activation Localization Operator), a lightweight white-box detector that operates on raw hidden-state volumes from a selected layer window. Across Qwen, Llama, and Mistral models, SALO improves jailbreak detection on several attack families under a fixed XSTest-calibrated operating point. We further analyze static RepE-style baselines, ROI sensitivity, adaptive GCG attacks, and encoded-input boundary cases, clarifying both the promise and limitations of refusal-trajectory monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。