攻击能跨模型通用,源于模型共享的表示空间。
Jailbreak Transferability Emerges from Shared Representations
- 通过良性提示微调增强模型表示相似性,可提升攻击迁移能力。
- 20个开源模型中,97%的攻击在高相似度模型间成功迁移。
- 自然语言类攻击比编码类攻击更易迁移,因前者利用共享语义空间。
Jailbreak迁移性指针对一个模型的对抗攻击能同时引发其他模型的有害响应,这一现象令人意外。尽管广泛观察到,但其成因尚无共识:是安全训练的偶然缺陷,还是模型家族特性,抑或表征学习的深层属性?我们提出证据表明,迁移性源于共享表征而非偶然缺陷。在20个开源模型与33种越狱攻击上,我们发现两个系统性影响因素:(1)良性提示下的表征相似性,(2)攻击在源模型上的强度。为超越相关性,我们通过仅用良性数据蒸馏刻意增强表征相似性,结果因果性地提升了迁移能力。定性分析显示不同攻击类型存在系统性迁移模式:角色型攻击迁移率远高于编码类提示,支持自然语言攻击利用共享语义空间,而编码类依赖特定漏洞且难以泛化。综上,越狱迁移性应被理解为表征对齐的结果,而非安全训练的脆弱副产物。
原文摘要 · Abstract (English)
Jailbreak transferability is the surprising phenomenon when an adversarial attack compromising one model also elicits harmful responses from other models. Despite widespread demonstrations, there is little consensus on why transfer is possible: is it a quirk of safety training, an artifact of model families, or a more fundamental property of representation learning? We present evidence that transferability emerges from shared representations rather than incidental flaws. Across 20 open-weight models and 33 jailbreak attacks, we find two factors that systematically shape transfer: (1) representational similarity under benign prompts, and (2) the strength of the jailbreak on the source model. To move beyond correlation, we show that deliberately increasing similarity through benign only distillation causally increases transfer. Our qualitative analyses reveal systematic transferability patterns across different types of jailbreaks. For example, persona-style jailbreaks transfer far more often than cipher-based prompts, consistent with the idea that natural-language attacks exploit models' shared representation space, whereas cipher-based attacks rely on idiosyncratic quirks that do not generalize. Together, these results reframe jailbreak transfer as a consequence of representation alignment rather than a fragile byproduct of safety training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。