用工具库优化提示词,让越狱攻击更高效
JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization

- 构建原子越狱提示工具库,通过统一优化生成更强攻击提示
- 在多个模型上提升成功率,同时减少成功所需攻击次数
- 适合研究大模型安全漏洞与对抗性提示的学者
越狱攻击揭示了大型语言模型(LLMs)持续存在的安全缺陷。现有无状态单轮方法存在权衡:手工构造的提示词表达力强但静态,而迭代提示优化虽可自适应,却常依赖低级变异,需大量目标查询。我们提出 JailbreakOPT,一个工具辅助的迭代单轮越狱提示优化框架。该框架将多样化的原子越狱提示组织成攻击工具库,并通过统一的回合内优化抽象进行组合,生成更强的独立攻击提示。为复用跨轮次经验,JailbreakOPT 将工具选择建模为上下文相关老虎机问题,采用上下文汤普森采样指导探索与利用。在多个目标 LLM 和攻击目标上的实验表明,JailbreakOPT 在提升攻击成功率(ASR)的同时,显著降低达成成功所需的攻击次数(No.A),优于原子单轮攻击及现有迭代优化基线。
原文摘要 · Abstract (English)
Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many target queries. We propose JailbreakOPT, a tool-assisted framework for improving iterative single-turn jailbreak prompt optimization. JailbreakOPT organizes diverse atomic jailbreak prompts into an attack tool library and composes them through a unified intra-episode optimization abstraction to generate stronger standalone attack prompts. To reuse experience across attack episodes, JailbreakOPT further frames tool selection as a contextual bandit problem and applies contextual Thompson sampling to guide exploration and exploitation based on past outcomes. Experiments across multiple target LLMs and attack goals show that JailbreakOPT improves attack success rate (ASR) while reducing the number of attacks until success (No.A) compared with atomic single-turn attacks and existing iterative optimization baselines. This paper may contain offensive or harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。