首个端到端网页数据科学评测,挑战AI完成真实复杂任务。
WebDS: An End-to-End Benchmark for Web-based Data Science
- 设计870个跨29个网站的多步骤网页任务,涵盖数据获取与分析全流程。
- 顶尖大模型仅完成15%任务,远低于人类90%准确率,暴露严重能力缺陷。
- 聚焦真实场景中的工具使用与信息整合,适合评估智能体实战能力。
许多现实中的数据科学任务涉及复杂的网络交互:在互联网上查找合适的数据、从不同来源整合多模态数据,并生成总结性分析。现有网页基准大多关注简单交互,缺乏多样工具使用要求;传统数据科学基准则集中于静态结构化数据集,未评估涵盖数据获取、清洗、分析和洞察生成的端到端工作流。为此,我们提出WebDS,首个端到端网页数据科学基准。它包含870个跨29个不同网站(从结构化政府数据门户到非结构化新闻媒体)的真实任务,要求智能体执行复杂、多步骤、基于工具的操作,处理异构数据格式,更真实反映现代数据分析的全貌。对当前SOTA大模型智能体的评估显示显著性能差距:如在WebVoyager上完成80%任务的Browser Use,在WebDS中仅完成15%。分析表明失败源于信息定位不准、重复行为及捷径策略。相比之下,人类达到约90%准确率,凸显当前智能体与人类表现的巨大差距。WebDS为发展实用的大模型数据科学能力提供了更可靠测试平台。
原文摘要 · Abstract (English)
Many real-world data science tasks involve complex web-based interactions: finding appropriate data available on the internet, synthesizing multimodal data from different locations, and producing summarized analyses. Existing web benchmarks often focus on simplistic interactions and often do not require diverse tool-using capabilities. Conversely, traditional data science benchmarks typically concentrate on static, highly structured datasets and do not assess end-to-end workflows that encompass data acquisition, cleaning, analysis, and insight generation. In response, we introduce WebDS, the first end-to-end web-based data science benchmark. It comprises 870 web-based data science tasks across 29 diverse websites from structured government data portals to unstructured news media, challenging agents to perform complex, multi-step, tool-based operations, across heterogeneous data formats, to better reflect the realities of modern data analytics. Evaluations of current SOTA LLM agents indicate significant performance gaps in accomplishing these tasks. For instance, Browser Use, which accomplishes $80\%$ of tasks on WebVoyager, completes only 15% of tasks in WebDS, which our analysis suggests is due to new failure modes, such as poor information grounding, repetitive behavior and shortcut-taking that agents performing WebDS's tasks display. By contrast, humans achieve around 90% accuracy, highlighting a substantial gap between current agents and human performance. By providing a more robust and realistic testing ground, WebDS sets the stage for significant advances in the development of practically useful LLM-based data science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。