arXiv:2511.00203cs.LGstat.ML2025-11被引 7

用扩散模型生成高效攻击提示,一次生成多个高危指令。

Diffusion LLMs are Natural Adversaries for any LLM

  • 用非自回归扩散语言模型直接生成攻击性提示,替代传统耗时优化。
  • 少量样本即可生成低困惑度、高转移性的越狱提示,效果显著。
  • 适合安全测试、红队对抗和自动化提示优化,尤其适用于新模型。

我们提出一种新框架,将资源密集型的对抗性提示优化问题转化为高效的可复用推理任务。核心思路是利用预训练的非自回归生成式语言模型(如扩散语言模型),其能建模提示-响应对的联合分布,作为提示搜索的强大代理。该方法可直接条件生成提示,用少量并行采样替代逐实例的离散优化。概率分析表明,在温和保真度假设下,仅需少数条件样本即可恢复高奖励(有害)提示。实证发现,生成的提示具有低困惑度、多样性高,且在多种黑盒目标模型(包括鲁棒训练和专有模型)上表现出强迁移能力。除对抗提示外,本框架还为红队测试、自动化提示优化及基于流与扩散的语言模型应用开辟了新路径。

原文摘要 · Abstract (English)

We introduce a novel framework that transforms the resource-intensive (adversarial) prompt optimization problem into an \emph{efficient, amortized inference task}. Our core insight is that pretrained, non-autoregressive generative LLMs, such as Diffusion LLMs, which model the joint distribution over prompt-response pairs, can serve as powerful surrogates for prompt search. This approach enables the direct conditional generation of prompts, effectively replacing costly, per-instance discrete optimization with a small number of parallelizable samples. We provide a probabilistic analysis demonstrating that under mild fidelity assumptions, only a few conditional samples are required to recover high-reward (harmful) prompts. Empirically, we find that the generated prompts are low-perplexity, diverse jailbreaks that exhibit strong transferability to a wide range of black-box target models, including robustly trained and proprietary LLMs. Beyond adversarial prompting, our framework opens new directions for red teaming, automated prompt optimization, and leveraging emerging Flow- and Diffusion-based LLMs.

对抗攻击扩散模型提示工程红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。