通过操控注意力机制,提升大模型越狱攻击成功率并降低生成成本。
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
- 利用注意力损失优化提示,动态调节模型关注点。
- 使GCG等攻击成功率提升至91.2%(原为67.9%),生成时间不足三分之一。
- 攻击具备跨模型迁移性,适用于多种越狱方法。
近期研究显示,精心设计的越狱输入可诱使大语言模型生成有害内容,即使经过对齐安全措施。为有效防御并准确评估模型安全性,需预判潜在越狱攻击范围。本文提出一种新方法,通过操纵模型注意力,选择性增强或抑制提示中不同部分的关注度。基于注意力损失,我们开发出更高效的越狱攻击,且具备可迁移性。该方法显著提升现有越狱算法(如GCG、AutoDAN、ReNeLLM)的成功率,同时降低生成成本:例如,在Llama2-7B/AdvBench上,增强版GCG攻击成功率达91.2%,远超原版的67.9%,生成时间不足原方法的三分之一。
原文摘要 · Abstract (English)
Recent research has shown that carefully crafted jailbreak inputs can induce large language models to produce harmful outputs, despite safety measures such as alignment. It is important to anticipate the range of potential Jailbreak attacks to guide effective defenses and accurate assessment of model safety. In this paper, we present a new approach for generating highly effective Jailbreak attacks that manipulate the attention of the model to selectively strengthen or weaken attention among different parts of the prompt. By harnessing attention loss, we develop more effective jailbreak attacks, that are also transferrable. The attacks amplify the success rate of existing Jailbreak algorithms including GCG, AutoDAN, and ReNeLLM, while lowering their generation cost (for example, the amplified GCG attack achieves 91.2% ASR, vs. 67.9% for the original attack on Llama2-7B/AdvBench, using less than a third of the generation time).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。