研究如何绕过大模型安全防护,实测成功率最高达100%。
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
- 用字符注入和对抗机器学习方法绕过检测
- 在6个主流系统上实现最高100%的绕过成功率
- 适合安全研究人员和模型防护开发者参考
大型语言模型(LLMs)的防护系统旨在抵御提示注入和越狱攻击,但仍易受规避技术影响。本文通过传统字符注入与算法级对抗机器学习(AML) evasion 技术,演示了绕过主流防护系统的方法。测试涵盖微软 Azure Prompt Shield 和元宇宙 Prompt Guard 等六款系统,结果显示两种方法均能有效规避检测,同时保持攻击有效性,在部分情况下实现高达100%的攻击成功规避率(ASR)。此外,研究发现攻击者可通过离线白盒模型计算词重要性,提升对黑盒目标的攻击成功率。结果揭示了当前大模型防护机制的漏洞,强调需构建更鲁棒的安全防线。
原文摘要 · Abstract (English)
Large Language Models (LLMs) guardrail systems are designed to protect against prompt injection and jailbreak attacks. However, they remain vulnerable to evasion techniques. We demonstrate two approaches for bypassing LLM prompt injection and jailbreak detection systems via traditional character injection methods and algorithmic Adversarial Machine Learning (AML) evasion techniques. Through testing against six prominent protection systems, including Microsoft's Azure Prompt Shield and Meta's Prompt Guard, we show that both methods can be used to evade detection while maintaining adversarial utility achieving in some instances up to 100% evasion success. Furthermore, we demonstrate that adversaries can enhance Attack Success Rates (ASR) against black-box targets by leveraging word importance ranking computed by offline white-box models. Our findings reveal vulnerabilities within current LLM protection mechanisms and highlight the need for more robust guardrail systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。