提出可插拔的多轮攻击框架,显著提升对大模型的越狱成功率。
PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits
- 分三阶段设计多轮攻击:引导、规划、收尾,模拟持续学习过程。
- 在少查询下攻击成功率超30%提升,o3模型越狱率达81.4%。
- 适合安全评估者研究模型漏洞,尤其关注多轮交互风险。
大型语言模型(LLMs)发展迅速,随着智能体工作流的出现,多轮对话已成为完成复杂任务的主流交互方式。然而,模型在多轮场景中更易受越狱攻击,有害意图可通过逐步渗透对话实现恶意输出。尽管单轮攻击已有广泛研究,但多轮攻击仍面临适应性差、效率低、效果弱等挑战。为此,本文提出PLAGUE框架,受终身学习智能体启发,将多轮攻击生命周期划分为三个精心设计阶段:引导(Primer)、规划(Planner)和收尾(Finisher),实现系统化且信息丰富的多轮攻击探索。实验表明,基于PLAGUE设计的红队智能体在多个主流模型上达到顶尖越狱效果,在更低或相当的查询预算下,攻击成功率(ASR)提升超30%。特别地,其在OpenAI o3模型上的强拒绝标准(StrongReject)ASR达81.4%,在Claude Opus 4.1上达67.3%,两者均为安全文献中公认的高抗越狱能力模型。本工作为理解计划初始化、上下文优化与终身学习在多轮攻击中的作用提供了工具与洞见,有助于全面评估模型脆弱性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are improving at an exceptional rate. With the advent of agentic workflows, multi-turn dialogue has become the de facto mode of interaction with LLMs for completing long and complex tasks. While LLM capabilities continue to improve, they remain increasingly susceptible to jailbreaking, especially in multi-turn scenarios where harmful intent can be subtly injected across the conversation to produce nefarious outcomes. While single-turn attacks have been extensively explored, adaptability, efficiency and effectiveness continue to remain key challenges for their multi-turn counterparts. To address these gaps, we present PLAGUE, a novel plug-and-play framework for designing multi-turn attacks inspired by lifelong-learning agents. PLAGUE dissects the lifetime of a multi-turn attack into three carefully designed phases (Primer, Planner and Finisher) that enable a systematic and information-rich exploration of the multi-turn attack family. Evaluations show that red-teaming agents designed using PLAGUE achieve state-of-the-art jailbreaking results, improving attack success rates (ASR) by more than 30% across leading models in a lesser or comparable query budget. Particularly, PLAGUE enables an ASR (based on StrongReject) of 81.4% on OpenAI's o3 and 67.3% on Claude's Opus 4.1, two models that are considered highly resistant to jailbreaks in safety literature. Our work offers tools and insights to understand the importance of plan initialization, context optimization and lifelong learning in crafting multi-turn attacks for a comprehensive model vulnerability evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。