无需标注数据,通过模拟越狱激活实现零样本防御。
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
- 无监督发现潜在方向,扩展越狱激活空间覆盖。
- 对抗训练使攻击成功率降至5%以下,通用性更强。
- 适合应对未知新式越狱攻击,提升大模型安全性。
越狱提示可诱导对齐的大语言模型生成有害内容。为此提出了安全引导机制:在测试时干预激活状态,将越狱激活导向拒绝响应,同时保持良性功能。然而现有方法依赖有监督且固定训练集,面对不断演化的、分布外的新越狱攻击时易失效。本文提出一种基于无监督潜在方向发现的双层对抗训练框架,实现零样本越狱防御。内层通过无监督方法从拒绝态有害请求激活中外推,模拟多样化越狱激活,扩展真实越狱激活子空间覆盖;外层训练一个势能诱导的引导场,将这些对抗性越狱状态推向拒绝区域,同时保持良性状态不变。在三个大模型和六类经典越狱攻击上,本方法攻击成功率普遍低于5%,且随训练过程子空间覆盖率持续上升,解释了其泛化能力提升。
原文摘要 · Abstract (English)
Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activations to trigger refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out-of-distributed from the training set, leading to failures on unseen attacks. In this paper, we tackle the failure on unseen jailbreaks problem, base on unsupervised latent direction discovery. We propose a bi-level adversarial training framework for zero-shot jailbreak defense. In the inner step, we simulate diverse jail-broken activations by extrapolating from refusal-state harmful-request activations via unsupervised latent direction discovery, which expands the coverage of real jailbreak activation subspaces. In the outer step, we train a potential-induced steering field to push these adversarial jailbroken states into refusal regions while keeping benign unchanged. Across three LLMs and six classical jailbreak families, our method achieves strong defense with attack success rates mostly below 5%, and rising subspace coverage throughout training helps explain the improved generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。