arXiv:2602.10481cs.CRcs.AI2026-02被引 2

用密码学方法确保提示和上下文可信,让大模型安全从被动防御变主动保障。

Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI

  • 用可验证的提示溯源和防篡改哈希链保障输入可信
  • 在6类攻击中实现100%检测率且零误报,开销极小
  • 适合需要高安全性大模型应用的金融、医疗等场景

大型语言模型应用易受提示注入和上下文操纵攻击,传统安全模型无法防范。本文提出两种新原语——认证提示与认证上下文,实现跨大模型工作流的密码学可验证来源追溯。认证提示支持自包含的路径验证,认证上下文则利用防篡改哈希链确保动态输入完整性。基于这些原语,我们形式化了一种策略代数,并证明了四个定理,实现协议级拜占庭容错——即使恶意代理也无法违反组织策略。五种互补防御机制(从轻量资源控制到大模型语义验证)构建分层预防性安全体系,具备形式化保证。对6类代表性攻击的全面评估显示,检测率达100%,零误报,计算开销极低。这是首个结合密码学强制提示溯源、防篡改上下文与可证明策略推理的方法,推动大模型安全从被动检测转向主动防护。

原文摘要 · Abstract (English)

Large Language Model (LLM) applications are vulnerable to prompt injection and context manipulation attacks that traditional security models cannot prevent. We introduce two novel primitives--authenticated prompts and authenticated context--that provide cryptographically verifiable provenance across LLM workflows. Authenticated prompts enable self-contained lineage verification, while authenticated context uses tamper-evident hash chains to ensure integrity of dynamic inputs. Building on these primitives, we formalize a policy algebra with four proven theorems providing protocol-level Byzantine resistance--even adversarial agents cannot violate organizational policies. Five complementary defenses--from lightweight resource controls to LLM-based semantic validation--deliver layered, preventative security with formal guarantees. Evaluation against representative attacks spanning 6 exhaustive categories achieves 100% detection with zero false positives and nominal overhead. We demonstrate the first approach combining cryptographically enforced prompt lineage, tamper-evident context, and provable policy reasoning--shifting LLM security from reactive detection to preventative guarantees.

大模型安全密码学提示保护可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。