arXiv:2605.26999cs.CLcs.CR2026-05

检测提示注入需考虑部署环境,不同场景下效果差异大。

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals

论文配图:Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals
图 1 · 摘自论文原文
  • 构建多模型多场景框架,评估提示注入检测在真实部署中的表现。
  • 结构信号可识别角色重定义等攻击模式,提升低误报率下的稳定性。
  • 无单一模型通用最优,需根据实际场景选择检测策略。

提示注入对大语言模型的安全部署构成重大威胁,但现有检测方法通常在有限设置下评估,无法反映真实运行约束。本文提出一种面向部署的评估框架,涵盖多模型、多运行模式实验,对比词汇、语义、结构及基于Transformer的检测器在多种分布外场景、重复数据划分下,以及排名与阈值化部署指标的表现。引入可解释的结构信号,捕捉层级覆盖、系统提示欺骗、角色重定义和规避模式,并评估其在稀疏模型及强编码基线中的贡献。结果表明,检测性能高度依赖运行模式且对阈值敏感,无模型在所有场景中占优;基于Transformer的模型整体表现最强,结构信号虽增益有限但在特定场景中稳定提升,尤其改善了高难度场景下的低误报表现。研究揭示了排名性能与实际部署有效性之间的差距,强调必须在真实操作条件下评估提示注入防御措施。代码将开源。

原文摘要 · Abstract (English)

Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated under limited settings that do not reflect real-world operating constraints. In this work, we present a deployment-aware evaluation of prompt injection detection using a multi-model and multi-regime experimental framework. We compare lexical, semantic, structural, and transformer-based detectors across multiple out-of-distribution settings, repeated data splits, and both ranking and thresholded deployment metrics. We introduce interpretable structural signals that capture hierarchy overrides, system prompt spoofing, role redefinition, and evasion patterns, and assess their contribution both within sparse models and in combination with strong encoder baselines. Our results show that detection performance is highly regime-dependent and sensitive to threshold selection, with no single model dominating across all settings. Transformer-based models achieve the strongest overall performance, while structural signals provide modest but consistent gains in certain regimes and improve low false positive rate behaviour in harder scenarios. These findings highlight the gap between ranking performance and deployment effectiveness and underscore the importance of evaluating prompt injection defences under realistic operational constraints. Code will be released.

提示注入模型安全部署评估结构信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。