测试AI能否在真实网站完成日常网络任务,发现现有模型表现有限。
ClawBench: Can AI Agents Complete Everyday Online Tasks?

- 构建153项真实网页任务的评估框架,覆盖15类生活工作场景。
- 8个前沿模型平均仅完成33.3%任务,主流模型仍难应对复杂流程。
- 基于生产环境动态网站,适合评估真实世界通用智能助手能力。
AI代理或许能处理邮件和文档,但能否可靠完成真实网站上的日常网络流程?日常生活任务为评估下一代AI代理提供了真实且未解决的测试环境。为此,我们提出ClawBench,一个包含153项人们日常生活中需定期完成的任务评估框架,涵盖144个平台、15个类别,从购物下单、预约挂号到提交求职申请。这些任务要求超越现有基准的能力,如从用户提供的文档中获取信息、跨多样化平台执行多步骤流程,以及大量文本填写等操作。与以往在离线沙箱中使用静态页面评估不同,ClawBench运行于生产级网站,保留真实网络环境的全部复杂性、动态性和交互挑战。通过拦截层捕获并阻止最终提交请求,确保评估安全且无真实副作用。对8个前沿模型的评估显示,无论是专有还是开源模型,仅完成极小部分任务。例如,Claude Sonnet 4.6仅达成33.3%,暴露出当前AI代理的重大差距。在ClawBench上的进展正推动我们向具备通用服务能力的AI助手迈进。
原文摘要 · Abstract (English)
AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework comprising 153 everyday online tasks that people need to accomplish regularly in their lives and work, spanning 144 platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and interaction challenges of real-world web environments. An interception layer captures and blocks the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 8 frontier models show that both proprietary and open-source models complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%, which exposes gaps in current AI agents. Progress on ClawBench brings us closer to AI agents that can function as general-purpose assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。