arXiv:2603.11619cs.CRcs.AI2026-03被引 27

分析自主大模型代理的安全风险并提出系统性防护框架

Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats

  • 构建五层生命周期安全框架,覆盖从初始化到执行的全过程
  • 发现提示注入、记忆污染等威胁在真实场景中普遍且严重
  • 强调需从点状防御转向整体架构防护,适合安全研究者参考

自主大语言模型(LLM)代理,以 OpenClaw 为例,展现出执行复杂长周期任务的卓越能力。然而,其紧密耦合的即时通讯交互模式和高权限执行特性显著扩大了系统攻击面。本文对 OpenClaw 进行全面安全威胁分析,提出一个面向生命周期的五层安全框架,涵盖初始化、输入、推理、决策与执行阶段,系统性地考察跨阶段复合威胁,包括间接提示注入、技能供应链污染、记忆污染及意图漂移。通过详细案例研究,揭示这些威胁在实际中的普遍性和严重性,并分析现有防御措施的局限性。结果表明,当前基于单点的防御机制难以应对跨时间、多阶段的系统性风险,凸显了构建整体安全架构的必要性。在此框架下,我们进一步评估各阶段代表性防御策略,如插件审核体系、上下文感知指令过滤、记忆完整性验证协议、意图验证机制与能力约束架构。

原文摘要 · Abstract (English)

Autonomous Large Language Model (LLM) agents, exemplified by OpenClaw, demonstrate remarkable capabilities in executing complex, long-horizon tasks. However, their tightly coupled instant-messaging interaction paradigm and high-privilege execution capabilities substantially expand the system attack surface. In this paper, we present a comprehensive security threat analysis of OpenClaw. To structure our analysis, we introduce a five-layer lifecycle-oriented security framework that captures key stages of agent operation, i.e., initialization, input, inference, decision, and execution, and systematically examine compound threats across the agent's operational lifecycle, including indirect prompt injection, skill supply chain contamination, memory poisoning, and intent drift. Through detailed case studies on OpenClaw, we demonstrate the prevalence and severity of these threats and analyze the limitations of existing defenses. Our findings reveal critical weaknesses in current point-based defense mechanisms when addressing cross-temporal and multi-stage systemic risks, highlighting the need for holistic security architectures for autonomous LLM agents. Within this framework, we further examine representative defense strategies at each lifecycle stage, including plugin vetting frameworks, context-aware instruction filtering, memory integrity validation protocols, intent verification mechanisms, and capability enforcement architectures.

大模型安全自主代理威胁分析防御框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。