arXiv:2603.12230cs.LGcs.AI2026-03被引 7

剖析前沿AI代理的新型安全风险与防御策略

Security Considerations for Artificial Intelligence Agents

  • 从代码与数据分离等基础假设出发,分析代理架构带来的新攻击面
  • 识别出间接提示注入、混乱代理人行为等关键安全漏洞
  • 适合关注AI系统安全设计的研究者与企业安全团队参考

本文基于Perplexity在数百万用户及数千家企业中运营通用型智能体的经验,探讨前沿AI代理的安全问题。智能体架构颠覆了代码与数据分离、权限边界和执行可预测性等传统假设,引发新的机密性、完整性与可用性失效模式。文章梳理了工具、连接器、托管边界及多智能体协同中的主要攻击面,重点关注间接提示注入、混乱代理人行为以及长流程工作流中的级联失败。随后评估了多层防御体系:输入层与模型层缓解措施、沙箱执行,以及对高后果操作的确定性策略控制。最后指出现有标准与研究的缺口,包括自适应安全基准、委托与权限控制的策略模型,以及符合NIST风险管理原则的多智能体系统安全设计指南。

原文摘要 · Abstract (English)

This article, a lightly adapted version of Perplexity's response to NIST/CAISI Request for Information 2025-0035, details our observations and recommendations concerning the security of frontier AI agents. These insights are informed by Perplexity's experience operating general-purpose agentic systems used by millions of users and thousands of enterprises in both controlled and open-world environments. Agent architectures change core assumptions around code-data separation, authority boundaries, and execution predictability, creating new confidentiality, integrity, and availability failure modes. We map principal attack surfaces across tools, connectors, hosting boundaries, and multi-agent coordination, with particular emphasis on indirect prompt injection, confused-deputy behavior, and cascading failures in long-running workflows. We then assess current defenses as a layered stack: input-level and model-level mitigations, sandboxed execution, and deterministic policy enforcement for high-consequence actions. Finally, we identify standards and research gaps, including adaptive security benchmarks, policy models for delegation and privilege control, and guidance for secure multi-agent system design aligned with NIST risk management principles.

AI安全智能体风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。