针对小模型提示注入攻击,提出自动防御生成框架。
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
- 用思维链种子迭代优化防御提示
- 降低攻击成功率与误报率,提升检测能力
- 适合边缘设备部署的小型开源模型安全防护
在大语言模型快速演进的背景下,本文聚焦于提示注入攻击带来的重大安全风险,尤其关注小型开源模型(如 LLaMA 系列)。提出一种新型防御机制,可自动生成防御策略,并在一组基准攻击上系统评估其效果。实验表明,该方法显著缓解了目标劫持漏洞。研究贡献包括:(1) 评估现有基于提示的防御对最新攻击的应对能力;(2) 提出一种以思维链(Chain of Thoughts)为种子、通过迭代优化防御提示的新框架;(3) 显著提升对目标劫持攻击的检测性能。所提策略在降低攻击成功率与误检率的同时,有效识别目标劫持行为,为资源受限环境下小型开源模型的安全高效部署提供支持。
原文摘要 · Abstract (English)
In this fast-evolving area of LLMs, our paper discusses the significant security risk presented by prompt injection attacks. It focuses on small open-sourced models, specifically the LLaMA family of models. We introduce novel defense mechanisms capable of generating automatic defenses and systematically evaluate said generated defenses against a comprehensive set of benchmarked attacks. Thus, we empirically demonstrated the improvement proposed by our approach in mitigating goal-hijacking vulnerabilities in LLMs. Our work recognizes the increasing relevance of small open-sourced LLMs and their potential for broad deployments on edge devices, aligning with future trends in LLM applications. We contribute to the greater ecosystem of open-source LLMs and their security in the following: (1) assessing present prompt-based defenses against the latest attacks, (2) introducing a new framework using a seed defense (Chain Of Thoughts) to refine the defense prompts iteratively, and (3) showing significant improvements in detecting goal hijacking attacks. Out strategies significantly reduce the success rates of the attacks and false detection rates while at the same time effectively detecting goal-hijacking capabilities, paving the way for more secure and efficient deployments of small and open-source LLMs in resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。