剖析开源智能体的四大安全风险与防护策略
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures

- 构建分层安全框架,识别智能体推理、执行与交互中的漏洞
- 揭示技能污染、认知操控等新型攻击威胁,暴露高权限风险
- 适合关注AI系统安全与可信部署的研究者和开发者
大型语言模型驱动的自主智能体快速发展,催生了OpenClaw这一新型开源代理框架。这类系统具备持续运行、技能增强、持久记忆和多通道交互能力,可自主完成复杂多步任务并无缝对接外部应用,但同时也带来了显著扩大的攻击面。尤其在高权限操作与持久内存的结合下,OpenClaw代理面临技能污染、认知操控、多智能体级联失效及供应链漏洞等新兴威胁。本文对OpenClaw代理的安全现状进行系统性综述:首先分析其架构特征与传统智能体的区别;其次建立分层威胁分类框架,解析推理、行动执行与外部交互中漏洞的成因;再回顾代表性防御机制,描绘当前防御图景;最后探讨生态可靠性与可信性方面的未解问题。
原文摘要 · Abstract (English)
The rapid evolution of large language model (LLM)-driven autonomous agents has given rise to OpenClaw, a new class of open-source agent frameworks that operate as continuously running, skill-augmented systems with persistent memory, multi-channel interaction, and high degrees of autonomy. Such capabilities enable OpenClaw agents to autonomously execute complex, multi-step tasks and interact seamlessly with external applications, but simultaneously introduce a substantially enlarged attack surface. In particular, the combination of high-privilege operations and persistent memory exposes OpenClaw agents to various emerging threats, including skill poisoning, cognitive manipulation, multi-agent cascading failures, and supply-chain vulnerabilities. In this survey, we present a comprehensive study of the security landscape of OpenClaw agents. We first examine the general architecture and key characteristics that distinguish OpenClaw agents from traditional AI agent systems. We categorize existing security and privacy threats into a layered framework and analyze how vulnerabilities arise during agent reasoning, action execution, and external interaction. Representative defense mechanisms are also reviewed to draw the current defense landscape. Finally, several unresolved issues related to the reliability and trustworthiness of OpenClaw ecosystems are discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。