arXiv:2510.17006cs.CL2025-10被引 4

通过在线优化提示词,动态防御大模型的迭代越狱攻击。

Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization

  • 用强化学习动态优化提示词,识别并拒绝有害请求。
  • 在3个大模型上对5种越狱攻击防御成功率超现有方法。
  • 兼顾安全与正常任务质量,适合高风险场景部署。

反复重写并输入提示词以诱导大语言模型产生有害输出的迭代越狱攻击,已被证明是极具威胁的攻击方式。现有防御方法无法主动打断这一动态试错过程。本文提出一种基于在线学习的新型防御框架,能针对每轮新输入动态更新策略。利用有害越狱提示与正常提示的差异,采用强化学习优化提示,确保对无害任务响应恰当,同时明确拒绝有害请求。为防止对攻击中部分提示重写过度拟合,引入过去方向梯度衰减(PDGD)。在三个大模型上测试显示,本方法显著优于五种现有防御策略,对五种迭代越狱攻击均表现更优,且同时提升了无害任务的响应质量。

原文摘要 · Abstract (English)

Iterative jailbreak methods that repeatedly rewrite and input prompts into large language models (LLMs) to induce harmful outputs -- using the model's previous responses to guide each new iteration -- have been found to be a highly effective attack strategy. Despite being an effective attack strategy against LLMs and their safety mechanisms, existing defenses do not proactively disrupt this dynamic trial-and-error cycle. In this study, we propose a novel framework that dynamically updates its defense strategy through online learning in response to each new prompt from iterative jailbreak methods. Leveraging the distinctions between harmful jailbreak-generated prompts and typical harmless prompts, we introduce a reinforcement learning-based approach that optimizes prompts to ensure appropriate responses for harmless tasks while explicitly rejecting harmful prompts. Additionally, to curb overfitting to the narrow band of partial input rewrites explored during an attack, we introduce Past-Direction Gradient Damping (PDGD). Experiments conducted on three LLMs show that our approach significantly outperforms five existing defense methods against five iterative jailbreak methods. Moreover, our results indicate that our prompt optimization strategy simultaneously enhances response quality for harmless tasks.

大模型安全在线学习提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。