arXiv:2604.06436cs.CRcs.AI2026-04被引 3

证明了提示注入防御封装无法完全安全,三者不可兼得。

The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?

  • 连续性、保持效用、完全防护三者无法同时成立
  • 任何防御在输入边界处必有失效区域
  • 适用于主流大模型,对防御设计具根本性指导

我们证明,在提示空间连通的语言模型中,不存在连续且保持效用的输入预处理防御函数(D: X→X),能确保所有输出严格安全。在逐步强化的假设下,我们得出三个结果:边界固定性——某些阈值输入必须不变;ε-鲁棒约束——在利普希茨条件下,边界附近存在正测度的近阈值区域;横截性条件下的持续不安全区域——存在正测度的输入子集始终不安全。这构成了防御三难困境:连续性、效用保留与完备性不可共存。我们还给出了无需拓扑的离散版本,并扩展至多轮交互、随机防御和容量对等情形。理论在Lean 4中机械验证,且在三个LLM上实证有效。研究不否定训练阶段对齐、架构改进或牺牲效用的防御方式。

原文摘要 · Abstract (English)

We prove that no continuous, utility-preserving wrapper defense-a function $D: X\to X$ that preprocesses inputs before the model sees them-can make all outputs strictly safe for a language model with connected prompt space, and we characterize exactly where every such defense must fail. We establish three results under successively stronger hypotheses: boundary fixation-the defense must leave some threshold-level inputs unchanged; an $ε$-robust constraint-under Lipschitz regularity, a positive-measure band around fixed boundary points remains near-threshold; and a persistent unsafe region under a transversality condition, a positive-measure subset of inputs remains strictly unsafe. These constitute a defense trilemma: continuity, utility preservation, and completeness cannot coexist. We prove parallel discrete results requiring no topology, and extend to multi-turn interactions, stochastic defenses, and capacity-parity settings. The results do not preclude training-time alignment, architectural changes, or defenses that sacrifice utility. The full theory is mechanically verified in Lean 4 and validated empirically on three LLMs.

提示注入防御三难大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。