arXiv:2409.19521cs.CRcs.LG2024-09被引 12

构建首个统一评测框架与防护系统,有效识别8万多条提示注入攻击

GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks

论文配图:GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
图 1 · 摘自论文原文
  • 提出新型检测模型GenTel-Shield,基于多维度特征分析攻击提示
  • 在28种场景中检测成功率超现有方法,暴露传统防护机制严重缺陷
  • 开源84812条攻击样本,适合安全研究者和大模型开发者使用

大型语言模型如GPT-4、LLaMA和Qwen在诸多应用中表现卓越,但仍易受提示注入攻击影响,此类攻击可绕过现有安全机制,凸显亟需更可靠的检测方法与全面的评估基准。为此,我们提出GenTel-Safe统一框架,包含新型提示注入攻击检测方法GenTel-Shield,以及涵盖84812条攻击样本的综合性评估基准GenTel-Bench,覆盖3大类别和28种安全场景。为验证GenTel-Shield有效性,我们将其与原始安全防护措施在GenTel-Bench上进行对比测试。实验表明,GenTel-Shield达到当前最优攻击检测成功率,揭示现有防护技术对恶意提示存在显著弱点。代码与数据集已公开于项目主页 https://gentellab.github.io/gentel-safe.github.io/。

原文摘要 · Abstract (English)

Large Language Models (LLMs) like GPT-4, LLaMA, and Qwen have demonstrated remarkable success across a wide range of applications. However, these models remain inherently vulnerable to prompt injection attacks, which can bypass existing safety mechanisms, highlighting the urgent need for more robust attack detection methods and comprehensive evaluation benchmarks. To address these challenges, we introduce GenTel-Safe, a unified framework that includes a novel prompt injection attack detection method, GenTel-Shield, along with a comprehensive evaluation benchmark, GenTel-Bench, which compromises 84812 prompt injection attacks, spanning 3 major categories and 28 security scenarios. To prove the effectiveness of GenTel-Shield, we evaluate it together with vanilla safety guardrails against the GenTel-Bench dataset. Empirically, GenTel-Shield can achieve state-of-the-art attack detection success rates, which reveals the critical weakness of existing safeguarding techniques against harmful prompts. For reproducibility, we have made the code and benchmarking dataset available on the project page at https://gentellab.github.io/gentel-safe.github.io/.

提示注入安全检测大模型防御评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。