通过模拟人类推理,识别并防御隐藏在提示中的恶意操作。
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning
- 用结构化推理链分析提示,发现隐藏的恶意操作模式。
- 在多个数据集上达到当前最优防御效果,对未知攻击泛化能力强。
- 适合关注大模型安全与对抗攻击防御的研究者和开发者。
防御大型语言模型(LLMs)免受越狱攻击对其实现安全可靠部署至关重要。现有防御方法多依赖浅层模式匹配,难以泛化到新型或未见过的攻击策略。为此,我们提出认知驱动防御(CDD)框架,通过应用元操作(meta-operations,即基本伪装有害意图的操作)来针对越狱提示的内在结构。CDD 通过结构化推理链模拟人类认知过程:先进行全局感知,再进行局部分析以发现隐藏操纵。通过对该推理链进行监督微调,模型学会识别已知操纵模式。为增强对未知威胁的泛化能力,引入熵引导的强化学习算法(EG-GRPO),鼓励探索新类型及变体的元操作。实验表明,CDD 能实现最先进的防御性能,并表现出对未见越狱攻击的强大泛化能力。
原文摘要 · Abstract (English)
Defending large language models (LLMs) against jailbreak attacks is essential for their safe and reliable deployment. Existing defenses often rely on shallow pattern matching, which struggles to generalize to novel and unseen attack strategies. To address this challenge, we propose the Cognitive-Driven Defense (CDD) framework, which targets the underlying structure of jailbreak prompts by applying meta-operations, defined as basic manipulations that conceal harmful intent.CDD emulates human cognitive reasoning through a structured reasoning chain. It begins with a global perception of the prompt and follows with a localized analysis to uncover hidden manipulations. By applying supervised fine-tuning on this structured chain, the model learns to identify and reason about known manipulation patterns. To enhance generalization to unseen threats, an entropy-guided reinforcement learning algorithm (EG-GRPO) is introduced to encourage exploration of new types and variants of meta-operations. Experiments demonstrate that CDD can achieve state-of-the-art defense performance and exhibit strong generalization to unseen jailbreak attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。