首次实测主流个人AI代理安全漏洞,发现其核心架构存在严重风险。
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
- 提出CIK三维度框架,从能力、身份、知识分析代理安全
- 单维度污染使攻击成功率从24.6%升至64%-74%,最安全模型仍超3倍
- 现有防御仍无效,文件保护虽阻97%攻击但影响正常更新
OpenClaw是2026年初部署最广泛的个人AI代理,具备本地系统全权限并集成Gmail、Stripe及文件系统等敏感服务。尽管高权限带来强大自动化能力,但也暴露出传统沙箱评估无法覆盖的安全风险。本文首次开展真实世界下的OpenClaw安全评估,提出CIK分类法,将代理的持久状态统一为能力(Capability)、身份(Identity)和知识(Knowledge)三个维度。评估涵盖12种攻击场景,在四种主干模型(Claude Sonnet 4.5、Opus 4.6、Gemini 3.1 Pro、GPT-5.4)上进行。结果表明,任一CIK维度被污染后,平均攻击成功率从24.6%升至64%-74%,即使最鲁棒模型也呈现三倍以上漏洞增长。我们进一步测试三种对齐CIK的防御策略与文件保护机制;最强防御在能力目标攻击下仍达63.8%成功率,文件保护可阻断97%恶意注入,但同时妨碍合法更新。整体表明,漏洞根植于代理架构,亟需更系统的防护机制。项目页面:https://ucsc-vlaa.github.io/CIK-Bench。
原文摘要 · Abstract (English)
OpenClaw, the most widely deployed personal AI agent in early 2026, operates with full local system access and integrates with sensitive services such as Gmail, Stripe, and the filesystem. While these broad privileges enable high levels of automation and powerful personalization, they also expose a substantial attack surface that existing sandboxed evaluations fail to capture. To address this gap, we present the first real-world safety evaluation of OpenClaw and introduce the CIK taxonomy, which unifies an agent's persistent state into three dimensions, i.e., Capability, Identity, and Knowledge, for safety analysis. Our evaluations cover 12 attack scenarios on a live OpenClaw instance across four backbone models (Claude Sonnet 4.5, Opus 4.6, Gemini 3.1 Pro, and GPT-5.4). The results show that poisoning any single CIK dimension increases the average attack success rate from 24.6% to 64-74%, with even the most robust model exhibiting more than a threefold increase over its baseline vulnerability. We further assess three CIK-aligned defense strategies alongside a file-protection mechanism; however, the strongest defense still yields a 63.8% success rate under Capability-targeted attacks, while file protection blocks 97% of malicious injections but also prevents legitimate updates. Taken together, these findings show that the vulnerabilities are inherent to the agent architecture, necessitating more systematic safeguards to secure personal AI agents. Our project page is https://ucsc-vlaa.github.io/CIK-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。