构建统一平台评估提示注入攻击,揭示现有防御的局限性。
PIArena: A Platform for Prompt Injection Evaluation

- 设计可扩展平台,支持多种攻击与防御的集成评测
- 发现主流防御在跨任务泛化上表现有限,易受自适应攻击突破
- 适合安全研究者和模型开发者快速验证防御方案
提示注入攻击在众多实际应用中带来严重安全风险。尽管关注度日益提升,但社区仍面临关键缺口:缺乏统一的提示注入评估平台,导致难以可靠比较防御方法、理解其在多样化攻击下的真实鲁棒性,或评估其跨任务和基准的泛化能力。例如,许多最初报告有效的防御在多样数据集和攻击下被发现鲁棒性有限。为此,我们提出PIArena,一个统一且可扩展的提示注入评估平台,支持用户轻松集成前沿攻击与防御,并在多种现有及新基准上进行评测。我们还设计了一种基于策略的动态攻击,可根据防御反馈自适应优化注入提示。通过PIArena的全面评估,我们揭示了当前先进防御的关键局限:跨任务泛化能力差、对自适应攻击脆弱,以及当注入任务与目标任务一致时的根本挑战。代码与数据集已公开于https://github.com/sleeepeer/PIArena。
原文摘要 · Abstract (English)
Prompt injection attacks pose serious security risks across a wide range of real-world applications. While receiving increasing attention, the community faces a critical gap: the lack of a unified platform for prompt injection evaluation. This makes it challenging to reliably compare defenses, understand their true robustness under diverse attacks, or assess how well they generalize across tasks and benchmarks. For instance, many defenses initially reported as effective were later found to exhibit limited robustness on diverse datasets and attacks. To bridge this gap, we introduce PIArena, a unified and extensible platform for prompt injection evaluation that enables users to easily integrate state-of-the-art attacks and defenses and evaluate them across a variety of existing and new benchmarks. We also design a dynamic strategy-based attack that adaptively optimizes injected prompts based on defense feedback. Through comprehensive evaluation using PIArena, we uncover critical limitations of state-of-the-art defenses: limited generalizability across tasks, vulnerability to adaptive attacks, and fundamental challenges when an injected task aligns with the target task. The code and datasets are available at https://github.com/sleeepeer/PIArena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。