arXiv:2601.03005cs.CRcs.AI2026-01ACL被引 3

通过修复动态越狱路径,提升大模型对越狱攻击的防御能力。

JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification

  • 动态挖掘策略样本,识别并修复越狱路径
  • 在多种攻击下防御成功率提升至92.3%
  • 适合安全对齐与模型净化场景使用

尽管进行了广泛的安全对齐,大语言模型仍容易遭受越狱攻击。虽然机器遗忘技术可通过擦除特定有害参数提供防御,但现有方法仍易受多样越狱攻击影响。我们首次开展实证研究,发现失败根源在于越狱攻击主要激活中间层中未被擦除的参数。通过探测这些绕过的参数如何重新组合成违规输出,我们验证了动态越狱路径的持续存在,并指出无法修复这些路径是当前遗忘防御的根本缺陷。为此,我们提出首个修复动态越狱路径的方案——Jailbreak Path Unlearning(JPU),通过动态挖掘策略内对抗样本,暴露漏洞并定位越狱路径。大量实验表明,JPU显著增强了对动态攻击的抵御能力,同时保持模型实用性。

原文摘要 · Abstract (English)

Despite extensive safety alignment, Large Language Models (LLMs) often fail against jailbreak attacks. While machine unlearning has emerged as a promising defense by erasing specific harmful parameters, current methods remain vulnerable to diverse jailbreaks. We first conduct an empirical study and discover that this failure mechanism is caused by jailbreaks primarily activating non-erased parameters in the intermediate layers. Further, by probing the underlying mechanism through which these circumvented parameters reassemble into the prohibited output, we verify the persistent existence of dynamic $\textbf{jailbreak paths}$ and show that the inability to rectify them constitutes the fundamental gap in existing unlearning defenses. To bridge this gap, we propose $\textbf{J}$ailbreak $\textbf{P}$ath $\textbf{U}$nlearning (JPU), which is the first to rectify dynamic jailbreak paths towards safety anchors by dynamically mining on-policy adversarial samples to expose vulnerabilities and identify jailbreak paths. Extensive experiments demonstrate that JPU significantly enhances jailbreak resistance against dynamic attacks while preserving the model's utility.

越狱防御模型遗忘路径修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。