arXiv:2507.19227cs.CL2025-07被引 22

发现扩散模型文本生成有严重安全漏洞,攻击成功率高达97%。

Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation

  • 提出并行解码攻击(PAD),利用多点注意力引导生成有害内容。
  • 在4个扩散模型上实现97%攻击成功率,危害生成速度提升2倍。
  • 揭示扩散架构与传统模型的安全差异,适合安全研究者参考。

大型语言扩散模型(LLDMs)在推理速度和数学推理任务中表现媲美大语言模型(LLMs),但其精确快速的生成能力加剧了有害内容生成的风险。现有针对LLMs的越狱方法对LLDMs效果有限,无法有效暴露其安全缺陷。防御措施虽存在,但尚不清楚LLDMs是否具备安全鲁棒性,或攻击本身不适用于扩散架构。为此,我们首次揭示了LLDMs的越狱脆弱性,证明攻击失败源于架构本质差异。提出并行解码越狱方法(PAD),引入多点注意力攻击,模仿LLMs中肯定响应模式,引导并行生成过程产生有害输出。在4个LLDM上实验显示,PAD攻击成功率高达97%,暴露出显著安全漏洞。此外,与同规模自回归LLMs相比,LLDMs的危害生成速度提升2倍,凸显未受控滥用的重大风险。通过全面分析,我们深入探究了LLDM架构特性,为扩散型语言模型的安全部署提供关键洞见。

原文摘要 · Abstract (English)

Large Language Diffusion Models (LLDMs) exhibit comparable performance to LLMs while offering distinct advantages in inference speed and mathematical reasoning tasks.The precise and rapid generation capabilities of LLDMs amplify concerns of harmful generations, while existing jailbreak methodologies designed for Large Language Models (LLMs) prove limited effectiveness against LLDMs and fail to expose safety vulnerabilities.Successful defense cannot definitively resolve harmful generation concerns, as it remains unclear whether LLDMs possess safety robustness or existing attacks are incompatible with diffusion-based architectures.To address this, we first reveal the vulnerability of LLDMs to jailbreak and demonstrate that attack failure in LLDMs stems from fundamental architectural differences.We present a PArallel Decoding jailbreak (PAD) for diffusion-based language models. PAD introduces Multi-Point Attention Attack, which guides parallel generative processes toward harmful outputs that inspired by affirmative response patterns in LLMs. Experimental evaluations across four LLDMs demonstrate that PAD achieves jailbreak attack success rates by 97%, revealing significant safety vulnerabilities. Furthermore, compared to autoregressive LLMs of the same size, LLDMs increase the harmful generation speed by 2x, significantly highlighting risks of uncontrolled misuse.Through comprehensive analysis, we provide an investigation into LLDM architecture, offering critical insights for the secure deployment of diffusion-based language models.

扩散模型越狱攻击安全漏洞LLDM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。