arXiv:2605.18672cs.AI2026-05被引 1

安全部署大模型智能体需三层概率保障架构

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

  • 分层设计:每层负责语义意图、环境有效性、动态可行性之一
  • 三层独立验证,通过概率链式法则保证整体安全
  • 适合关注大模型智能体安全落地的研究者与工程师

本文主张,在单一抽象层中强制大模型智能体安全不仅低效,更是结构性不足——这是由智能体执行过程本质决定的,而非当前系统局限。安全运行依赖三个维度:语义意图与策略合规性、环境有效性、动态可行性,它们分别依赖执行不同阶段才可获取的独立信息。单一防护机制无法覆盖全部维度。我们提出基于合约的三层数学架构,每一层独立认证,其概率保证作为下一层的前提。利用概率链式法则推导出系统级安全边界。当前仍存在三大挑战:从非独立同分布轨迹估计边界、部署漂移下的合约优雅退化、多智能体场景扩展——这些是大模型智能体运行保障最核心的未解决问题。

原文摘要 · Abstract (English)

This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for deployed LLM agents -- a structural consequence of how agent execution works, not a contingent limitation of current systems. The three dimensions that jointly constitute safe operation -- semantic intent and policy compliance, environmental validity, and dynamical feasibility -- each depend on a strictly distinct set of information that becomes available at different stages of execution. No single guardrail can certify all three. We argue that the community must respond with a contract-based architecture in which each safety dimension is enforced by an independently certified layer whose probabilistic guarantee satisfies the next layer's assumption. We sketch such an architecture and derive the compositional system-level safety bounds it admits via the chain rule of probability. Three open problems stand between this and a deployable standard: bound estimation from non-i.i.d.\ traces, graceful degradation of contracts under deployment drift, and extension to multi-agent settings -- the most important unfinished business in LLM agent runtime assurance.

大模型安全智能体概率保障系统架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。