评估自托管工具型AI在6类风险下的安全表现,发现模糊指令易引发严重误操作。
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
- 以交互轨迹为中心,结合自动评分与人工审查评估安全风险。
- 34个案例中多数失败源于意图不明确或看似无害的越狱提示。
- 适合关注AI代理安全、工具调用风险的研究者与开发者参考。
Clawdbot 是一个自托管的、可使用多种工具的个人AI代理,其行动空间涵盖本地执行与网络协作流程,因此在存在歧义和恶意引导时面临更高的安全与隐私风险。本文基于六类风险维度,开展以轨迹为中心的评估。测试集从先前的代理安全基准(如ATBench和LPS-Bench)中采样并轻度适配,同时补充了针对Clawdbot工具表面设计的手动案例。完整记录交互轨迹(消息、动作、工具调用参数与输出),并使用自动化轨迹裁判器(AgentDoG-Qwen3-4B)与人工评审进行安全评估。在34个典型案例中,系统在可靠性任务上表现稳定,但多数失败出现在意图不明确、开放目标或看似无害的越狱提示场景下,微小误解可能演变为高影响的工具操作。通过代表性案例分析,总结了常见漏洞与失效模式,揭示了Clawdbot在实际应用中易触发的安全问题。
原文摘要 · Abstract (English)
Clawdbot is a self-hosted, tool-using personal AI agent with a broad action space spanning local execution and web-mediated workflows, which raises heightened safety and security concerns under ambiguity and adversarial steering. We present a trajectory-centric evaluation of Clawdbot across six risk dimensions. Our test suite samples and lightly adapts scenarios from prior agent-safety benchmarks (including ATBench and LPS-Bench) and supplements them with hand-designed cases tailored to Clawdbot's tool surface. We log complete interaction trajectories (messages, actions, tool-call arguments/outputs) and assess safety using both an automated trajectory judge (AgentDoG-Qwen3-4B) and human review. Across 34 canonical cases, we find a non-uniform safety profile: performance is generally consistent on reliability-focused tasks, while most failures arise under underspecified intent, open-ended goals, or benign-seeming jailbreak prompts, where minor misinterpretations can escalate into higher-impact tool actions. We supplemented the overall results with representative case studies and summarized the commonalities of these cases, analyzing the security vulnerabilities and typical failure modes that Clawdbot is prone to trigger in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。