构建更全面的个人助手评测基准,检验模型在复杂数字环境中的持续服务能力。
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

- 扩展三维度上下文:长期行为历史、多服务联动、跨设备交互界面。
- GPT-5.5在新基准上仅34.5%通过率,远低于旧标准。
- 适合评估智能助手的主动服务与抗干扰能力,推动真实场景应用。
大型语言模型代理正被构想为能随时访问用户数字世界全部信息的个人助理。然而当前系统仅能接触有限信息片段,限制了上下文感知推理和有效协助。现有评测也仅提供部分用户状态,无法反映这种广域、持续运行场景下的表现。为此,我们提出Claw-Anything基准,从三个维度拓展代理上下文:长周期活动历史、相互依赖的后端服务,以及跨多个设备的图形与命令行界面整合。通过多轮事件注入模拟数月用户行为,生成复杂世界状态和真实噪声(如无关事件与矛盾信号),要求代理在丰富上下文中推理并保持鲁棒性。该设定还支持主动服务评估,要求代理预判用户需求并及时推荐。实验显示,GPT-5.5仅获34.5% pass@1,显著低于以往基准,凸显当前代理能力与始终在线个人助理需求间的差距。同时,我们发布自动化数据生成管道,生成2000个训练环境,使基础模型性能提升23.7%,验证了可扩展数据基础设施的价值。
原文摘要 · Abstract (English)
Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。