arXiv:2505.15753cs.CRcs.AI2025-05被引 9

用检索增强生成技术提升大模型对抗越狱攻击的防御能力

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

  • 通过检索安全对齐样本增强模型抗性
  • 在多种越狱攻击下实现更优防御效果
  • 适合关注大模型安全落地的研究者

大型语言模型易受越狱攻击,攻击者通过精心设计的提示诱导模型产生有害或不道德回应。尽管已有防御机制部分缓解风险,但新型对抗技术不断突破现有防护,暴露静态防御框架的局限。本文基于上下文检索视角探索应对演化中的越狱威胁:首先发现仅需少量安全对齐示例即可显著提升对特定攻击模式的鲁棒性;在此基础上,结合检索增强生成(RAG)技术,提出安全上下文检索(SCR)机制,构建可扩展、强鲁棒的模型防护范式。大量实验表明,SCR在应对已知与新兴越狱策略时均表现优异,为大模型安全提供新思路。代码将在发表后公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are known to be vulnerable to jailbreaking attacks, wherein adversaries exploit carefully engineered prompts to induce harmful or unethical responses. Such threats have raised critical concerns about the safety and reliability of LLMs in real-world deployment. While existing defense mechanisms partially mitigate such risks, subsequent advancements in adversarial techniques have enabled novel jailbreaking methods to circumvent these protections, exposing the limitations of static defense frameworks. In this work, we explore defending against evolving jailbreaking threats through the lens of context retrieval. First, we conduct a preliminary study demonstrating that even a minimal set of safety-aligned examples against a particular jailbreak can significantly enhance robustness against this attack pattern. Building on this insight, we further leverage the retrieval-augmented generation (RAG) techniques and propose Safety Context Retrieval (SCR), a scalable and robust safeguarding paradigm for LLMs against jailbreaking. Our comprehensive experiments demonstrate how SCR achieves superior defensive performance against both established and emerging jailbreaking tactics, contributing a new paradigm to LLM safety. Our code will be available upon publication.

大模型安全越狱攻击RAG防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。