梳理自主智能体的新型安全威胁与防护方法
Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- 构建了面向自主智能体的威胁分类体系
- 综述了现有评估基准与防御策略
- 适合关注AI安全的开发者与政策制定者
由大语言模型驱动、具备规划、工具使用、记忆和自主性能力的自主智能体系统,正成为自动化的重要灵活平台。其在网页、软件及物理环境中自主执行任务的能力,带来了区别于传统AI安全和常规软件安全的新且被放大的安全风险。本文提出了一种针对自主智能体的威胁分类体系,综述了近期的基准测试与评估方法,并从技术和治理双重视角讨论了防御策略。通过整合当前研究并指出开放挑战,旨在推动安全设计的智能体系统发展。
原文摘要 · Abstract (English)
Agentic AI systems powered by large language models (LLMs) and endowed with planning, tool use, memory, and autonomy, are emerging as powerful, flexible platforms for automation. Their ability to autonomously execute tasks across web, software, and physical environments creates new and amplified security risks, distinct from both traditional AI safety and conventional software security. This survey outlines a taxonomy of threats specific to agentic AI, reviews recent benchmarks and evaluation methodologies, and discusses defense strategies from both technical and governance perspectives. We synthesize current research and highlight open challenges, aiming to support the development of secure-by-design agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。