arXiv:2602.08412cs.AI2026-02被引 31

为个性化AI助手设计安全评估框架,发现其多阶段漏洞。

From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent

  • 构建真实场景下的端到端安全评测框架PASB。
  • OpenClaw在提示、工具调用、记忆检索中均现严重漏洞。
  • 适合研究个性化AI安全与攻防的学者和工程师。

尽管基于大语言模型的智能体(如OpenClaw)正从任务导向系统演变为解决复杂现实任务的个性化助手,但其实际部署也带来了严峻的安全风险。现有安全研究与评估框架多聚焦于合成或任务中心场景,难以准确捕捉个性化智能体在真实部署中的攻击面与风险传播机制。为此,我们提出个性化智能体安全评测框架(PASB),一个面向真实场景的端到端安全评估体系。PASB融合个性化使用场景、真实工具链与长周期交互,支持对真实系统的黑盒、端到端安全评估。以OpenClaw为例,我们系统评估了其在多种个性化场景、工具能力与攻击类型下的安全性。结果表明,OpenClaw在用户提示处理、工具调用与记忆检索等不同执行阶段均存在关键漏洞,揭示了个性化智能体部署中的显著安全风险。相关代码已开源:https://github.com/AstorYH/PASB。

原文摘要 · Abstract (English)

Although large language model (LLM)-based agents, exemplified by OpenClaw, are increasingly evolving from task-oriented systems into personalized AI assistants for solving complex real-world tasks, their practical deployment also introduces severe security risks. However, existing agent security research and evaluation frameworks primarily focus on synthetic or task-centric settings, and thus fail to accurately capture the attack surface and risk propagation mechanisms of personalized agents in real-world deployments. To address this gap, we propose Personalized Agent Security Bench (PASB), an end-to-end security evaluation framework tailored for real-world personalized agents. Building upon existing agent attack paradigms, PASB incorporates personalized usage scenarios, realistic toolchains, and long-horizon interactions, enabling black-box, end-to-end security evaluation on real systems. Using OpenClaw as a representative case study, we systematically evaluate its security across multiple personalized scenarios, tool capabilities, and attack types. Our results indicate that OpenClaw exhibits critical vulnerabilities at different execution stages, including user prompt processing, tool usage, and memory retrieval, highlighting substantial security risks in personalized agent deployments. The code for the proposed PASB framework is available at https://github.com/AstorYH/PASB.

AI安全智能体漏洞评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。