发现对齐大模型仍可被越狱的根源:拒绝-回答方向及其结构成因。
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

- 从连续输入变换视角揭示越狱本质是拒绝行为的渐变转移。
- 证明越狱方向可分解为归一化、残差连接等算子层的贡献。
- 揭示安全与实用性的权衡困境,指导更稳健的对齐设计。
对齐的大语言模型仍易受越狱攻击。尽管已有研究识别出与越狱成功相关的潜在特征和表征变化,但更根本的问题仍未解答:为何对齐模型仍可被越狱?其模型内部存在何种结构性漏洞?本文通过连续输入变换视角展开研究。理论发现,对齐模型仍可能表现出拒绝-逃逸方向(Refusal-Escape Directions, RED):在有害输入附近存在的局部扰动方向,能在保持有害语义理解的同时,使模型行为从拒绝回答转向回应。这意味着越狱不仅是离散提示构造的成功,更是沿红色方向持续扰动有害输入引发的拒绝-回答行为转变。进一步证明,红色可精确分解为模型算子结构中各算子层级来源的贡献,并识别出归一化、残差连接和终端模块为解析受限的算子级来源。要消除红色,共享表达模块(自注意力与MLP)必须消除这些解析受限来源的贡献,同时保留支持良性响应的机制。这种相互冲突的要求引出了条件性的安全-效用权衡。多模型与攻击方法的实验从两个互补角度验证了红色的存在,显示增加词元维度会暴露红色,而成功的越狱攻击显示出与终端源贡献高度一致的拒绝-回答偏移。
原文摘要 · Abstract (English)
Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representation shifts associated with jailbreak success, but they leave a more fundamental question open: why do aligned LLMs remain jailbreakable, and what structural vulnerabilities in the model make this possible? We study this question through a continuous input-transformation view. Our theoretical finding is that aligned models can still exhibit Refusal-Escape Directions (RED): local perturbation directions around a harmful input that shift the model's behavior from refusal to answering while preserving the model's harmful-semantics interpretation. From this perspective, a jailbreak is not only a successful discrete prompt construction, but can also be understood as a refusal-to-answer behavior transition induced by continuously perturbing a harmful input along RED. We then prove that RED can be exactly decomposed into contributions from operator-level sources across the model's operator structure, and identify normalization, residual-wiring, and terminal sources as analytically constrained operator-level sources. To eliminate RED, the shared expressive modules -- self-attention and MLP -- must eliminate the contributions from these analytically constrained sources while preserving the mechanisms that support benign responses. These competing requirements give rise to a conditional safety-utility trade-off. Experiments across multiple models and attack methods empirically analyze RED from two complementary perspectives and show that added token dimensions can expose RED, while successful jailbreaks exhibit refusal-to-answer shifts largely aligned with terminal-source contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。