攻击者先动更危险,新攻击方法突破12种主流防御
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- 设计自适应攻击,动态优化策略绕过防御
- 对12个防御系统攻击成功率超90%,原报告接近零失败率
- 适合研究模型安全与对抗攻击的学者参考
当前对大语言模型防御机制的评估通常基于静态攻击样本或计算能力弱的优化方法,这些方法并未针对防御设计。我们提出应评估防御在面对自适应攻击者时的表现——即攻击者主动调整策略、投入大量资源优化攻击目标。通过系统性地调优和扩展通用优化技术(梯度下降、强化学习、随机搜索、人工引导探索),我们成功以超过90%的攻击成功率突破了12种近期提出的防御方案,而这些防御原先报告的攻击成功率接近零。我们认为,未来防御研究必须考虑此类强攻击,才能做出可靠且可信的鲁棒性声明。
原文摘要 · Abstract (English)
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful knowledge or remotely triggering malicious actions, respectively) are typically evaluated either against a static set of harmful attack strings, or against computationally weak optimization methods that were not designed with the defense in mind. We argue that this evaluation process is flawed. Instead, we should evaluate defenses against adaptive attackers who explicitly modify their attack strategy to counter a defense's design while spending considerable resources to optimize their objective. By systematically tuning and scaling general optimization techniques-gradient descent, reinforcement learning, random search, and human-guided exploration-we bypass 12 recent defenses (based on a diverse set of techniques) with attack success rate above 90% for most; importantly, the majority of defenses originally reported near-zero attack success rates. We believe that future defense work must consider stronger attacks, such as the ones we describe, in order to make reliable and convincing claims of robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。