提出分层自治安全框架,应对AI代理从思考到协作的新型威胁
From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
- 按认知、执行、群体三层次划分AI代理安全层级
- 识别出推理操控、环境干扰、多代理系统失效等关键威胁
- 适合研究可信AI系统与多智能体安全的学者参考
人工智能代理已从被动预测工具演变为具备自主决策与环境交互能力的主动实体,这主要得益于大语言模型(LLMs)的推理能力。然而,这一演进引入了现有安全框架无法覆盖的关键漏洞。本文提出分层自治演化(HAE)框架,将代理安全划分为三个层级:认知自治(L1)关注内部推理完整性;执行自治(L2)涵盖工具驱动的环境交互;集体自治(L3)处理多代理生态中的系统性风险。我们构建了涵盖认知操纵、物理环境破坏和多代理系统故障的威胁分类体系,并评估现有防御措施,识别出关键研究空白。研究旨在指导可信AI代理系统中多层次、自治感知的防御架构设计。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) agents have evolved from passive predictive tools into active entities capable of autonomous decision-making and environmental interaction, driven by the reasoning capabilities of Large Language Models (LLMs). However, this evolution has introduced critical security vulnerabilities that existing frameworks fail to address. The Hierarchical Autonomy Evolution (HAE) framework organizes agent security into three tiers: Cognitive Autonomy (L1) targets internal reasoning integrity; Execution Autonomy (L2) covers tool-mediated environmental interaction; Collective Autonomy (L3) addresses systemic risks in multi-agent ecosystems. We present a taxonomy of threats spanning cognitive manipulation, physical environment disruption, and multi-agent systemic failures, and evaluate existing defenses while identifying key research gaps. The findings aim to guide the development of multilayered, autonomy-aware defense architectures for trustworthy AI agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。