通过模拟人类学习,自动提炼攻击特征并自适应优化防御策略。
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
- 模仿人类试错过程,自主构建攻击模式图谱与防御框架
- 在常规和高隐蔽性测试集上均显著超越现有方法,误检率更低
- 适合需要持续迭代的现实场景,如内容安全与可信AI部署
大型语言模型虽在多领域表现卓越,但仍易受对抗性越狱攻击。现有提示防御方法(包括参数修改与无参数方法)在适应性、可解释性与定制化方面存在局限。为此,我们提出ShieldLearner,一种类人学习的新范式:通过试错自主提炼攻击特征形成模式图谱,并生成防御启发式规则的元分析框架,实现系统化、可解释的威胁检测。同时引入自适应对抗增强技术,对成功防御的提示生成对抗变体,实现无需重训练的持续自我优化。除标准基准外,我们还基于Wildjailbreak数据集构建了高隐蔽性测试集。实验表明,ShieldLearner在常规与硬测试集上均显著优于现有基线,且计算开销更低,具备实际应用价值。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable success in various domains but remain vulnerable to adversarial jailbreak attacks. Existing prompt-defense strategies, including parameter-modifying and parameter-free approaches, face limitations in adaptability, interpretability, and customization, constraining their effectiveness against evolving threats. To address these challenges, we propose ShieldLearner, a novel paradigm that mimics human learning in defense. Through trial and error, it autonomously distills attack signatures into a Pattern Atlas and synthesizes defense heuristics into a Meta-analysis Framework, enabling systematic and interpretable threat detection. Furthermore, we introduce Adaptive Adversarial Augmentation to generate adversarial variations of successfully defended prompts, enabling continuous self-improvement without model retraining. In addition to standard benchmarks, we create a hard test set by curating adversarial prompts from the Wildjailbreak dataset, emphasizing more concealed malicious intent. Experimental results show that ShieldLearner achieves a significantly higher defense success rate than existing baselines on both conventional and hard test sets, while also operating with lower computational overhead, making it a practical and efficient solution for real-world adversarial defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。