扩散语言模型生成更安全,但嵌套上下文可破解其防护。
Safer by Diffusion, Broken by Context: Diffusion LLM's Safety Blessing and Its Failure Mode
- 通过扩散轨迹逐步抑制有害输出,提升安全性。
- 嵌套上下文攻击使多模型攻击成功率达新高。
- 揭示扩散模型安全机制的漏洞,适合安全研究者关注。
扩散型大语言模型(D-LLMs)在生成效率上优于自回归模型(AR-LLMs),且我们发现其扩散生成方式具备未被充分探索的安全优势:对专为AR-LLMs设计的越狱攻击具有内在鲁棒性。本文分析其机制,表明扩散路径通过逐步抑制效应降低有害生成风险。然而该安全并非绝对;我们提出一种简单有效的失败模式——上下文嵌套,即在结构化良性上下文中嵌入有害请求。实验证明,该黑盒策略在多个模型与基准上实现当前最高攻击成功率,首次成功越狱已知的Gemini Diffusion模型,暴露出专有扩散模型的关键漏洞。本工作系统刻画了D-LLMs安全优势的来源与边界,构成对扩散模型的早期红队测试。
原文摘要 · Abstract (English)
Diffusion large language models (D-LLMs) offer an alternative to autoregressive LLMs (AR-LLMs) and have demonstrated advantages in generation efficiency. Beyond the utility benefits, we argue that D-LLMs exhibit a previously underexplored safety blessing: their diffusion-style generation confers intrinsic robustness against jailbreak attacks originally designed for AR-LLMs. In this work, we provide an initial analysis of the underlying mechanism, showing that the diffusion trajectory induces a stepwise reduction effect that progressively suppresses unsafe generations. This robustness, however, is not absolute. Following this analysis, we highlight a simple yet effective failure mode, context nesting, in which harmful requests are embedded within structured benign contexts. Empirically, we show that this simple black-box strategy bypasses D-LLMs' safety blessing, achieving state-of-the-art attack success rates across models and benchmarks. Notably, it enables the first successful jailbreak of Gemini Diffusion to our knowledge, exposing a critical vulnerability in proprietary D-LLMs. Together, our results characterize both the origins and the limits of D-LLMs' safety blessing, constituting an early-stage red-teaming of D-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。