用动态生成的镜像提示对抗各类越狱攻击,提升模型安全性。
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting

- 通过构建反映输入结构的动态镜像提示,实现安全约束。
- 在多个数据集上抵御十种主流越狱攻击,防御效果更优。
- 适合关注大模型安全、需要通用防御方案的研究者使用。
防御大语言模型(LLMs)免受越狱攻击对保障其安全部署至关重要。现有防御策略通常依赖预设的静态规则区分有害与良性提示,但此类僵化规则难以应对真实世界越狱攻击的复杂性与动态性。本文聚焦于应对多样化越狱攻击的通用防御新挑战,提出“镜像”新概念——一种动态生成的提示,既能反映输入的句法结构,又能保证语义安全。输入提示与其对应镜像之间的差异,成为防御的指导原则。据此提出新型防御模型 MirrorShield,基于生成的镜像检测并校准高风险输入。在多个基准数据集上评估,并与十种先进攻击方法对比,MirrorShield展现出卓越的防御性能和良好的泛化能力。
原文摘要 · Abstract (English)
Defending large language models (LLMs) against jailbreak attacks is crucial for ensuring their safe deployment. Existing defense strategies typically rely on predefined static criteria to differentiate between harmful and benign prompts. However, such rigid rules fail to accommodate the inherent complexity and dynamic nature of real-world jailbreak attacks. In this paper, we focus on the novel challenge of universal defense against diverse jailbreaks. We propose a new concept ``mirror'', which is a dynamically generated prompt that reflects the syntactic structure of the input while ensuring semantic safety. The discrepancies between input prompts and their corresponding mirrors serve as guiding principles for defense. A novel defense model, MirrorShield, is further proposed to detect and calibrate risky inputs based on the crafted mirrors. Evaluated on multiple benchmark datasets and compared against ten state-of-the-art attack methods, MirrorShield demonstrates superior defense performance and promising generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。