arXiv:2412.12621cs.CL2024-12ACL被引 15

一招搞定越狱攻击,无需重写策略即可跨模型生效

Jailbreaking? One Step Is Enough!

  • 将攻击意图伪装成防御,诱导模型自洽生成有害内容
  • 单次迭代即成功越狱,跨模型攻击无需重设计
  • 结合小样本提示学习,提升攻击隐蔽性与成功率

大型语言模型在多项任务中表现优异,但易受越狱攻击影响,攻击者通过操纵提示生成有害内容。现有方法与模型防御相互独立,需频繁迭代并针对不同模型重新设计。为此,我们提出逆向嵌入防御攻击(REDA)机制,将攻击意图伪装为‘防御’有害内容的正当行为。具体而言,REDA从目标响应出发,引导模型将有害内容嵌入其防御措施中,使有害内容退居次要角色,让目标模型误认为自身在执行防御任务。攻击模型以为自己在指导防御,而目标模型则真以为在做安全防护,形成协作假象。此外,为增强模型对‘防御’意图的信心与引导,我们采用少量攻击示例进行上下文学习,并构建对应的数据集。大量实验表明,REDA可实现跨模型越狱攻击,无需重设计攻击策略,单次迭代即可成功,且在开源与闭源模型上均优于现有方法。

原文摘要 · Abstract (English)

Large language models (LLMs) excel in various tasks but remain vulnerable to jailbreak attacks, where adversaries manipulate prompts to generate harmful outputs. Examining jailbreak prompts helps uncover the shortcomings of LLMs. However, current jailbreak methods and the target model's defenses are engaged in an independent and adversarial process, resulting in the need for frequent attack iterations and redesigning attacks for different models. To address these gaps, we propose a Reverse Embedded Defense Attack (REDA) mechanism that disguises the attack intention as the "defense". intention against harmful content. Specifically, REDA starts from the target response, guiding the model to embed harmful content within its defensive measures, thereby relegating harmful content to a secondary role and making the model believe it is performing a defensive task. The attacking model considers that it is guiding the target model to deal with harmful content, while the target model thinks it is performing a defensive task, creating an illusion of cooperation between the two. Additionally, to enhance the model's confidence and guidance in "defensive" intentions, we adopt in-context learning (ICL) with a small number of attack examples and construct a corresponding dataset of attack examples. Extensive evaluations demonstrate that the REDA method enables cross-model attacks without the need to redesign attack strategies for different models, enables successful jailbreak in one iteration, and outperforms existing methods on both open-source and closed-source models.

越狱攻击对抗攻击LLM安全单次攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。