arXiv:2605.29251cs.AIcs.CR2026-05

用逻辑约束让智能体行动可验证,防攻击零失败。

Provably Secure Agent Guardrail

论文配图:Provably Secure Agent Guardrail
图 1 · 摘自论文原文
  • 将意图转为一阶逻辑公式,强制形式化执行
  • 对抗测试中攻击成功率与误报率均为零
  • 适合高安全要求的自主系统开发

随着大语言模型从生成引擎演变为拥有广泛执行权限的智能体,人工智能失控问题引发安全危机。现有防御架构依赖经验性语义护栏和概率模型判别器,在面对复杂语义符号解耦攻击时无法提供确定性安全下界。为此,本文提出基于逻辑推理本质限制的安全新范式,并引入可执行的证明约束动作(ePCA)框架,采用神经符号隔离结构。该框架摒弃自然语言语义信任,要求智能体在执行物理操作前将意图无损形式化为一阶逻辑数学约束。宏观与微观二维动态对抗系统的实证评估表明,该形式化验证机制在所有测试场景中实现零攻击成功率与零误报率,计算延迟极低。本研究在明确系统假设下建立了条件形式基础,为未来智能系统构建底层防御提供了工程范式。

原文摘要 · Abstract (English)

As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates a fundamental crisis in artificial intelligence security. Existing defense architectures heavily rely on empirical semantic guardrails and probabilistic large model adjudicators, mechanisms that fail to provide deterministic security lower bounds when facing complex semantic symbol decoupling attacks. To overcome this empirical semantic guardrail dilemma, this paper proposes a new security paradigm for agents based on the fundamental limitations of logical reasoning. Based on this paradigm, we further introduce an executable Proof-Constrained Action (ePCA) framework with a neural symbolic isolation architecture. This framework abandons semantic trust in natural language, forcing agents to losslessly formalize their intentions into first-order logical mathematical constraints before performing physical operations. Empirical evaluations of macroscopic and microscopic two-dimensional dynamic adversarial systems demonstrate that our formal verification mechanism achieves zero attack success rate and zero false positive rate across the evaluated scenarios, with extremely low computational latency. This research provides a conditional formal foundation under explicit system assumptions and an engineering paradigm for constructing the underlying defense foundation for future intelligent systems.

智能体安全形式验证逻辑约束对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。