arXiv:2508.16406cs.CRcs.CL2025-08ACL被引 1

用已有攻击案例库提升大模型防御能力,无需重训就能应对新攻击。

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

  • 利用已知攻击案例检索匹配潜在恶意请求
  • 对强攻击如PAP、PAIR的失效率显著降低,误拒率低
  • 可调节安全与可用性平衡,适合实际部署

大型语言模型(LLMs)仍易受越狱攻击,此类攻击旨在诱使模型生成有害内容。攻击方式不断演变且多样化,给防御系统带来挑战,包括(1)在不需昂贵重训练的情况下适应新攻击策略,(2)控制安全与可用性之间的权衡。为此,我们提出检索增强防御(RAD),一种新型越狱检测框架,将已知攻击示例数据库融入检索增强生成,用于推断用户背后的恶意查询和越狱策略。RAD支持无训练更新以应对新发现的越狱方法,并提供安全与实用性之间权衡的调控机制。在StrongREJECT上的实验表明,RAD显著降低了PAP和PAIR等强越狱攻击的有效性,同时保持对良性查询的低拒绝率。我们提出一种新颖评估方案,证明RAD可在多种操作点上以可控方式实现稳健的安全-效用平衡。

原文摘要 · Abstract (English)

Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature and diversity of these attacks pose many challenges for defense systems, including (1) adaptation to counter emerging attack strategies without costly retraining, and (2) control of the trade-off between safety and utility. To address these challenges, we propose Retrieval-Augmented Defense (RAD), a novel framework for jailbreak detection that incorporates a database of known attack examples into Retrieval-Augmented Generation, which is used to infer the underlying, malicious user query and jailbreak strategy used to attack the system. RAD enables training-free updates for newly discovered jailbreak strategies and provides a mechanism to balance safety and utility. Experiments on StrongREJECT show that RAD substantially reduces the effectiveness of strong jailbreak attacks such as PAP and PAIR while maintaining low rejection rates for benign queries. We propose a novel evaluation scheme and show that RAD achieves a robust safety-utility trade-off across a range of operating points in a controllable manner.

越狱防御检索增强LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。