让智能体行为全程合规,防止看似合法的步骤组合出安全问题。
Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

- 从单步检查升级为全程轨迹验证,确保整体行为合规
- 揭示了多智能体协作中的信任与控制漏洞,尤其在任务委托时
- 适合关注AI安全、系统可靠性及合规部署的研究者和工程师
自主智能体在受操作约束、组织政策、监管要求和技术标准约束的环境中执行重要任务,其安全性不再取决于单个动作的正确性,而在于整体行为是否符合系统规则与安全不变量。随着基于大语言模型的智能体越来越自主,并跨组织边界进行任务委托,安全问题已从单一挑战演变为覆盖整个智能体架构栈的复杂网络。在单智能体层面,提示、记忆、检索知识和工具接口等未受信任输入构成攻击面;在多智能体场景中,委托与通信带来身份、信任、能力控制和决策透明度问题,底层模型路由与执行控制平面也易遭篡改和模型来源不可信。最根本挑战是行为可容性:单个允许的动作序列可能共同违反系统级约束。在更宏观层面,供应链完整性、溯源性、问责制与端到端可观测性仍属开放问题。统一原则是:安全必须成为架构、协议与运行时的可验证属性,而非可选的指导层。明确这些挑战为可信自主智能体部署提供了路线图。
原文摘要 · Abstract (English)
Autonomous agents are increasingly used to execute consequential tasks in environments governed by operational constraints, organizational policies, regulatory requirements, and technical standards. Their safety is therefore determined not by the correctness of individual actions, but by whether their overall behavior remains consistent with the rules and invariants of the systems in which they operate. As large language model (LLM)-based agents become more autonomous and increasingly delegate tasks across organizational boundaries, securing them evolves from a single challenge into a broad and interconnected landscape spanning the entire agentic stack. At the single-agent level, untrusted inputs through prompts, memory, retrieved knowledge, and tool interfaces create attack surfaces. In multi-agent settings, delegation and communication introduce challenges related to identity, trust, capability control, and decision transparency, while the underlying model routing and execution control plane remains vulnerable to manipulation and to unverified model provenance. Perhaps the most fundamental challenge is behavioral containment: sequences of individually permissible actions may collectively violate system-level constraints and safety invariants. At the broader level, supply-chain integrity, provenance, accountability, and end-to-end observability remain largely open problems. A common principle unifies these directions: security must become a verifiable property of the architectures, protocols, and runtimes that govern agent behavior, rather than an optional layer of guidance. Charting these challenges provides a roadmap toward trustworthy autonomous agent deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。