构建长上下文提示注入基准,揭示现有防御在真实场景中的严重漏洞。
LongPIBench: A Long-Context Benchmark for Prompt Injection

- 针对论文评审等4类真实场景设计长文本攻击数据集
- 长上下文下简单攻击成功率超90%,突破主流防御
- 适合安全研究者和大模型应用开发者参考
提示注入攻击对大语言模型在实际应用中的安全性构成严重威胁。然而,现有提示注入基准主要聚焦于短上下文输入,导致长上下文场景下的攻击与防御研究长期被忽视,进而高估了现有防御的有效性。本文提出LongPIBench,一个面向长上下文的提示注入基准,涵盖论文同行评审、简历筛选、代码审查和邮件摘要四种真实应用场景。每个场景均构建合成数据集与真实世界数据集,上下文长度达数千至数万token。在LongPIBench上的评估结果显示,在长上下文设置下,即使采用简单启发式攻击,成功率仍极高,且频繁绕过当前最先进的防御机制。我们期望LongPIBench能成为系统评估真实长上下文场景下提示注入防御能力的实用基准。
原文摘要 · Abstract (English)
Prompt injection attacks pose a serious security risk to large language models in real-world applications. However, existing prompt injection benchmarks primarily focus on short-context inputs, leaving the attacks and defenses in long-context settings largely unexplored. This gap leads to a substantial overestimation of the effectiveness of current defenses. In this paper, we bridge the gap by introducing LongPIBench, a long-context benchmark for prompt injection covering 4 realistic application scenarios: paper peer review, resume screening, code review, and email summary. For each scenario, we construct a synthetic dataset and a real-world dataset, with context lengths ranging from thousands to tens of thousands of tokens. The evaluation results on LongPIBench reveal significant vulnerabilities of prompt injection defenses under long-context settings: even simple heuristic prompt injection attacks achieve high success rates and frequently bypass state-of-the-art defenses. We hope LongPIBench can serve as a practical benchmark for systematically evaluating prompt injection defenses in realistic long-context scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。