首个覆盖多领域的真实数据科学任务评估基准,助力大模型数据代理能力评测
AgenticDataBench: A Comprehensive Benchmark for Data Agents

- 构建跨15个垂直领域的真实任务集,含金融等5个企业级场景
- 基于技能聚类生成高质任务,覆盖9大核心数据操作技能
- 支持细粒度技能级评估,适合研究数据智能自动化者使用
数据科学旨在从异构原始数据中提取可行动的洞察,释放现代社会海量数据的价值。自动化该过程对降低数据科学家的劳动强度、实现可扩展的数据驱动应用至关重要。近年来,基于大语言模型(LLM)的数据代理成为自动化数据科学工作流的有前景方案。然而,该领域缺乏能跨多样化场景、以细粒度进行严谨评估的综合性基准。为此,我们提出AgenticDataBench,一个包含真实任务、覆盖多元领域且具备细粒度标注的综合基准。首先,为覆盖多样领域,我们从15个垂直领域收集真实数据集与任务,包括来自头部金融科技公司的5个真实B2B应用场景。其次,为消除真实任务冗余并生成高质量任务,我们引入数据科学技能、常见数据操作模式,并以技能数量衡量基准覆盖度。代表性技能通过在Stack Overflow上大规模任务解法上采用技能对齐的层次聚类提取。第三,针对真实商业任务,我们选择技能构成多样性最大化的工作流-解决方案对,确保实际场景广泛覆盖。第四,针对缺乏真实任务的领域,我们提出系统化的基于LLM的任务生成方法,依据这些技能生成真实工作流与任务。最后,我们利用标注好的基准与开源测试平台评估当前最先进数据代理,提供详细的技能级别洞察。
原文摘要 · Abstract (English)
Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to reducing labor-intensive efforts for data scientists and enabling scalable data-driven applications. Recently, large language model (LLM)-based data agents have emerged as a promising solution to automate data science workflows. However, the field lacks comprehensive benchmarks to rigorously evaluate these agents across diverse scenarios with fine-grained granularity. To address this gap, we propose AgenticDataBench, a comprehensive benchmark featuring realistic tasks spanning diverse domains with fine-grained ground-truth labels. This enables evaluations to capture the diversity and complexity of data science workflows and the detailed performance of agents. First, to cover diverse domains, we collect real datasets and tasks from 15 vertical domains, including 5 real-world B2B use cases from a leading fintech company. Second, to remove redundancy in real-world tasks and generate high-quality tasks for domains lacking real data, we introduce data science skills, recurring data-centric operational patterns, and quantify benchmark coverage by the number of skills included. Representative skills are extracted from large-scale task solutions on Stack Overflow using skill-aligned hierarchical clustering. Third, for real-world business tasks, we select task-solution pairs that maximize diversity in skill composition, ensuring broad coverage of practical scenarios. Fourth, to generate realistic tasks for devise domains without real tasks, we propose a systematic LLM-based task generation approach to create workflows and tasks based on these skills. Finally, we evaluate state-of-the-art data agents using our annotated benchmark and open-sourced testbed, providing detailed skill-level insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。