提出新基准CAPTURE,评估大模型提示注入防御的实战能力
CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement
- 构建上下文感知的动态测试基准,避免静态攻击数据偏差
- 发现现有防御模型在恶意攻击中漏报率高、正常请求误报多
- 基于该基准训练的新模型显著降低两类错误,通用性更强
提示注入仍是大语言模型的主要安全风险。现有防护模型在上下文感知场景下的有效性尚未充分探索,因其多依赖静态攻击基准且易过度防御。本文提出CAPTURE,一种新型上下文感知基准,可在极少领域内样本下评估攻击检测能力与过度防御倾向。实验表明,当前提示注入防护模型在对抗性案例中存在高漏报率,在良性场景中则误报过多,暴露出关键缺陷。为验证框架价值,我们基于生成数据训练了CaptureGuard。该模型在上下文感知数据集上显著降低漏报与误报率,并有效泛化至外部基准,为构建更鲁棒、实用的提示注入防御提供了新路径。
原文摘要 · Abstract (English)
Prompt injection remains a major security risk for large language models. However, the efficacy of existing guardrail models in context-aware settings remains underexplored, as they often rely on static attack benchmarks. Additionally, they have over-defense tendencies. We introduce CAPTURE, a novel context-aware benchmark assessing both attack detection and over-defense tendencies with minimal in-domain examples. Our experiments reveal that current prompt injection guardrail models suffer from high false negatives in adversarial cases and excessive false positives in benign scenarios, highlighting critical limitations. To demonstrate our framework's utility, we train CaptureGuard on our generated data. This new model drastically reduces both false negative and false positive rates on our context-aware datasets while also generalizing effectively to external benchmarks, establishing a path toward more robust and practical prompt injection defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。