arXiv:2608.03499cs.AI2026-08

构建可审计的跨用户智能体协作测试平台,评估安全风险与协同效率。

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

论文配图:WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
图 1 · 摘自论文原文
  • 设计可验证的跨用户协作沙箱,模拟真实个人工作空间下的多代理交互。
  • 包含620种任务变体,每项基础任务有1个正常场景和4个攻击变体,覆盖6大领域。
  • 支持对隐私泄露、权限滥用等攻击行为的实时审计与归因分析,适合安全研究者使用。

当前持久化个人智能体框架正推动以人为中心的智能体网络走向实际部署:每位用户可由一个代表其行动、维持状态并与其他智能体通过社交和任务关系通信的AI代理服务。在此类网络中,日常工具使用变为跨所有者的工作空间协作,文件、记录、工具与策略在所有者之间不可见。现有智能体基准测试关注工具使用与协作,但缺乏端到端可验证的跨用户协作沙箱,也无法测试有害行为在人机智能体网络中的传播路径。本文提出WeClawArena,一个面向多主体拥有的个人工作空间协作的可审计基准与运行时沙箱。该平台聚焦于以个人工作空间作为操作工具与个人约束的协同任务。基准包含124个基础任务,分布在六个跨用户任务领域,并扩展为620种场景变体,每个基础任务对应1个良性控制场景和4个攻击向量变体。沙箱记录同行消息、工具调用、资源操作、受控决策及最终工作空间状态。WeClawArena分别报告任务效用与攻击成功率,并基于有限运行时证据审计攻击成功,支持对任务失败、隐私泄露、污染证据和无效权限路径的诊断。

原文摘要 · Abstract (English)

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.

智能体协作安全审计跨用户沙箱测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。