探索用GCG攻击扩散型语言模型,发现其存在可被利用的漏洞。
GCG Attack On A Diffusion LLM
- 将GCG攻击方法用于扩散式语言模型,测试不同扰动策略
- 在AdvBench数据集上成功诱导有害输出,验证攻击有效性
- 为扩散模型安全研究提供新方向,适合关注AI安全的研究者
尽管多数大语言模型采用自回归生成,基于扩散的模型最近成为一种替代方案。贪婪坐标梯度(GCG)攻击在自回归模型中表现有效,但其在扩散语言模型中的适用性仍不明确。本文对开源扩散语言模型LLaDA进行探索性研究,评估多种攻击变体,包括前缀扰动和基于后缀的对抗生成,针对来自AdvBench数据集的有害提示进行实验。研究初步揭示了扩散语言模型的鲁棒性与攻击面,推动该领域对抗分析优化与评估策略的发展。
原文摘要 · Abstract (English)
While most LLMs are autoregressive, diffusion-based LLMs have recently emerged as an alternative method for generation. Greedy Coordinate Gradient (GCG) attacks have proven effective against autoregressive models, but their applicability to diffusion language models remains largely unexplored. In this work, we present an exploratory study of GCG-style adversarial prompt attacks on LLaDA (Large Language Diffusion with mAsking), an open-source diffusion LLM. We evaluate multiple attack variants, including prefix perturbations and suffix-based adversarial generation, on harmful prompts drawn from the AdvBench dataset. Our study provides initial insights into the robustness and attack surface of diffusion language models and motivates the development of alternative optimization and evaluation strategies for adversarial analysis in this setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。