arXiv:2510.21910cs.LG2025-10被引 5

通过复用旧攻击技能,提升模型对未知越狱攻击的防御能力。

Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks

  • 将攻击技能提取为稀疏词典,用组合方式模拟新攻击
  • 在32篇论文中验证:新攻击可由旧技能稀疏组合解释
  • 新训练法ASCoT显著增强对未见攻击的鲁棒性

大型语言模型仍易受越狱攻击影响,这些攻击能绕过安全防护生成有害内容。防御新型越狱攻击是人工智能安全的关键挑战。传统对抗训练因优化困难和威胁建模不现实,难以应对新出现的越狱攻击。本文提出新范式,基于‘对抗性既视感’假设:新攻击并非全新,而是已有攻击技能的重组。通过对两年内32篇攻击论文的大规模分析,我们使用自动化管道提取并压缩攻击技能为稀疏词典,由大模型生成可读描述。分析显示,未见过的攻击可被早期技能的稀疏组合有效解释,解释力随技能覆盖度单调提升。据此提出对抗技能组合训练(ASCoT),在多样化的技能组合上训练,而非孤立攻击实例。ASCoT显著提升对未见攻击(包括多轮越狱)的鲁棒性,同时保持低过度拒绝率。还证明扩大对抗技能覆盖比单纯增加数据量更关键。

原文摘要 · Abstract (English)

Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Defending against novel jailbreaks represents a critical challenge in AI safety. Adversarial training -- designed to make models robust against worst-case perturbations -- has been the dominant paradigm for adversarial robustness. However, due to optimization challenges and difficulties in defining realistic threat models, adversarial training methods often fail on newly developed jailbreaks in practice. This paper proposes a new paradigm for improving robustness against unseen jailbreaks, centered on the Adversarial Déjà Vu hypothesis: novel jailbreaks are not fundamentally new, but largely recombinations of adversarial skills from previous attacks. We study this hypothesis through a large-scale analysis of 32 attack papers published over two years. Using an automated pipeline, we extract and compress adversarial skills into a sparse dictionary of primitives, with LLMs generating human-readable descriptions. Our analysis reveals that unseen attacks can be effectively explained as sparse compositions of earlier skills, with explanatory power increasing monotonically as skill coverage grows. Guided by this insight, we introduce Adversarial Skill Compositional Training (ASCoT), which trains on diverse compositions of skill primitives rather than isolated attack instances. ASCoT substantially improves robustness to unseen attacks, including multi-turn jailbreaks, while maintaining low over-refusal rates. We also demonstrate that expanding adversarial skill coverage, not just data scale, is key to defending against novel attacks. \textcolor{red}{\textbf{Warning: This paper contains content that may be harmful or offensive in nature.

越狱攻击对抗训练安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。