arXiv:2603.08234cs.AIcs.LG2026-03

揭示大模型狱卒攻击的内在机制:延续冲动与安全防御的对抗

The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs

  • 分析注意力头,发现模型延续倾向与安全防御存在内部竞争
  • 仅调整指令后缀位置,攻击成功率显著提升
  • 为提升模型安全性提供可解释的新思路,适合安全研究者

随着大语言模型(LLMs)的快速发展,其安全性成为关键问题。尽管已有大量安全对齐努力,当前模型仍易受越狱攻击。本文聚焦一种由延续触发的越狱现象:仅通过移动延续触发的指令后缀,即可显著提高越狱成功率。通过注意力头级别的机制可解释性分析,结合因果干预与激活缩放,我们发现该行为主要源于模型内在的延续驱动力与对齐训练中习得的安全防御之间的内在竞争。进一步分析发现,不同模型架构中的安全关键注意力头在功能和行为上存在显著差异。这些发现为理解大模型越狱行为提供了新的机制视角,兼具理论价值与实践意义。

原文摘要 · Abstract (English)

With the rapid advancement of large language models (LLMs), the safety of LLMs has become a critical concern. Despite significant efforts in safety alignment, current LLMs remain vulnerable to jailbreaking attacks. However, the root causes of such vulnerabilities are still poorly understood, necessitating a rigorous investigation into jailbreak mechanisms across both academic and industrial communities. In this work, we focus on a continuation-triggered jailbreak phenomenon, whereby simply relocating a continuation-triggered instruction suffix can substantially increase jailbreak success rates. To uncover the intrinsic mechanisms of this phenomenon, we conduct a comprehensive mechanistic interpretability analysis at the level of attention heads. Through causal interventions and activation scaling, we show that this jailbreak behavior primarily arises from an inherent competition between the model's intrinsic continuation drive and the safety defenses acquired through alignment training. Furthermore, we perform a detailed behavioral analysis of the identified safety-critical attention heads, revealing notable differences in the functions and behaviors of safety heads across different model architectures. These findings provide a novel mechanistic perspective for understanding and interpreting jailbreak behaviors in LLMs, offering both theoretical insights and practical implications for improving model safety.

模型安全越狱攻击可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。