arXiv:2604.27464cs.CRcs.AI2026-04综述

系统梳理智能代理框架的四层安全风险与防御策略。

Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study

论文配图:Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
图 1 · 摘自论文原文
  • 按上下文、工具、状态、生态四层分析安全风险
  • 发现攻击可跨层传播,从输入到系统级影响
  • 适合研究智能代理安全与可信系统的学者

基于大语言模型的自主代理框架正发展为复杂、工具集成且持续运行的系统,其安全风险已超越传统提示层漏洞。尽管已有研究关注代理系统的不同攻击面和防御问题,但现有工作仍分散于单一维度,缺乏系统性分层综述。本文提出针对自主代理框架的安全风险与防御策略的分层综述,以OpenClaw为案例,将分析划分为四层:上下文与指令层、工具与动作层、状态与持久化层、生态与自动化层。每层总结功能角色、典型安全风险及对应防御措施。研究表明,攻击可跨层传播,由被篡改输入引发不安全动作,导致状态污染,并造成生态级影响。最后指出关键挑战:各层研究失衡、缺乏长时评估机制、生态信任模型薄弱,并提出未来系统化整合防御的方向。

原文摘要 · Abstract (English)

Autonomous agent frameworks built upon large language models (LLMs) are evolving into complex, tool-integrated, and continuously operating systems, introducing security risks beyond traditional prompt-level vulnerabilities. As this paradigm is still at an early stage of development, a timely and systematic understanding of its security implications is increasingly important. Although a growing body of work has examined different attack surfaces and defense problems in agent systems, existing studies remain scattered across individual aspects of agent security, and there is still a lack of a layered review on this topic. To address this gap, this survey presents a layered review of security risks and defense strategies in autonomous agent frameworks, with OpenClaw as a case study. We organize the analysis into four security-relevant layers: the context and instruction layer, the tool and action layer, the state and persistence layer, and the ecosystem and automation layer. For each layer, we summarize its functional role, representative security risks, and corresponding defense strategies. Based on this layered analysis, we further identify that threats in autonomous agent frameworks may propagate across layers, from manipulated inputs to unsafe actions, persistent state contamination, and broader ecosystem-level impact. Finally, we highlight potential key challenges, including research imbalance across layers, the lack of long-horizon evaluation, and weak ecosystem trust models, and outline future directions toward more systematic and integrated defenses.

智能代理安全风险分层分析防御策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。