arXiv:2410.22143cs.CL2024-10被引 10

用生成模型高效制造胡言乱语后缀,轻松突破大模型安全限制。

AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts

  • 训练生成模型自动产出可定制的无意义对抗后缀。
  • 对Llama-2和GPT-4攻击成功率提升最高达17%,黑盒下超三倍。
  • 适用于研究模型安全漏洞,尤其关注防御机制的开发者。

尽管大型语言模型(LLMs)通常经过对齐,但仍可能通过精心设计的自然语言提示或看似无意义的对抗后缀被破解。其中,无意义标记虽受关注度较低,却在攻击对齐模型方面表现优异。近期工作AmpleGCG证明,生成模型可快速产生大量可定制的无意义对抗后缀,暴露出分布外(OOD)语言空间中的多种对齐漏洞。为推动该方向发展,本文提出AmpleGCG-Plus,其在更少尝试下实现更高性能。通过一系列探索性实验,我们识别出若干优化无意义后缀学习的训练策略。严格评估结果显示,该方法在开源与闭源模型上均优于AmpleGCG:白盒攻击中对Llama-2-7B-chat的攻击成功率达17%提升;黑盒攻击中对GPT-4的成功率超过三倍。值得注意的是,AmpleGCG-Plus对最新GPT-4o系列模型的破解效率与GPT-4相当,并揭示了新型‘电路断路器’防御机制的潜在漏洞。本文公开发布AmpleGCG-Plus及训练数据集。

原文摘要 · Abstract (English)

Although large language models (LLMs) are typically aligned, they remain vulnerable to jailbreaking through either carefully crafted prompts in natural language or, interestingly, gibberish adversarial suffixes. However, gibberish tokens have received relatively less attention despite their success in attacking aligned LLMs. Recent work, AmpleGCG~\citep{liao2024amplegcg}, demonstrates that a generative model can quickly produce numerous customizable gibberish adversarial suffixes for any harmful query, exposing a range of alignment gaps in out-of-distribution (OOD) language spaces. To bring more attention to this area, we introduce AmpleGCG-Plus, an enhanced version that achieves better performance in fewer attempts. Through a series of exploratory experiments, we identify several training strategies to improve the learning of gibberish suffixes. Our results, verified under a strict evaluation setting, show that it outperforms AmpleGCG on both open-weight and closed-source models, achieving increases in attack success rate (ASR) of up to 17\% in the white-box setting against Llama-2-7B-chat, and more than tripling ASR in the black-box setting against GPT-4. Notably, AmpleGCG-Plus jailbreaks the newer GPT-4o series of models at similar rates to GPT-4, and, uncovers vulnerabilities against the recently proposed circuit breakers defense. We publicly release AmpleGCG-Plus along with our collected training datasets.

对抗攻击大模型安全生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。