DAComp构建企业级数据智能全流程评测基准,揭示现有智能体在工程与分析上的双重短板。
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- 覆盖210个任务,涵盖从数据工程到分析的全生命周期流程
- 工程类任务成功率不足20%,分析类任务平均得分低于40%
- 采用多指标评估与可信LLM裁判,适合研发企业级自主数据代理
真实企业数据智能工作流包括数据工程(将原始数据转化为可分析表)和数据分析(将表格转化为决策洞察)。我们提出DAComp,一个包含210个任务的基准,模拟这些复杂流程。数据工程任务要求在工业级数据模式上进行仓库级工程,包括从零设计多阶段SQL流水线,并在需求变化下演化已有系统。数据分析任务则为开放式商业问题,需战略规划、迭代编码探索、中间结果解读及可操作建议整合。工程任务通过执行驱动的多指标评估,开放式任务由经实验验证的LLM裁判评分,其依据分层精细设计的评分标准。实验表明,即使最先进的智能体在DAComp上表现不佳。工程任务成功率低于20%,暴露出整体流水线编排的关键瓶颈,而非仅代码生成问题。分析任务平均分低于40%,凸显开放推理能力严重不足,证明工程与分析是不同能力。通过明确诊断这些局限,DAComp为推动企业级自主数据代理发展提供严谨真实的测试环境。数据与代码已公开于https://da-comp.github.io
原文摘要 · Abstract (English)
Real-world enterprise data intelligence workflows encompass data engineering that turns raw sources into analytical-ready tables and data analysis that convert those tables into decision-oriented insights. We introduce DAComp, a benchmark of 210 tasks that mirrors these complex workflows. Data engineering (DE) tasks require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data analysis (DA) tasks pose open-ended business problems that demand strategic planning, exploratory analysis through iterative coding, interpretation of intermediate results, and the synthesis of actionable recommendations. Engineering tasks are scored through execution-based, multi-metric evaluation. Open-ended tasks are assessed by a reliable, experimentally validated LLM-judge, which is guided by hierarchical, meticulously crafted rubrics. Our experiments reveal that even state-of-the-art agents falter on DAComp. Performance on DE tasks is particularly low, with success rates under 20%, exposing a critical bottleneck in holistic pipeline orchestration, not merely code generation. Scores on DA tasks also average below 40%, highlighting profound deficiencies in open-ended reasoning and demonstrating that engineering and analysis are distinct capabilities. By clearly diagnosing these limitations, DAComp provides a rigorous and realistic testbed to drive the development of truly capable autonomous data agents for enterprise settings. Our data and code are available at https://da-comp.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。