构建真实数据科学任务评估基准,揭示现有智能体能力不足。
DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
- 设计包含466个分析与74个建模任务的真实场景基准
- 顶尖智能体仅完成34.12%分析任务,相对差距达34.74%
- 适合研究智能数据科学家、多模态推理与自动化建模的学者
大型语言模型(LLMs)和大型视觉语言模型(LVLMs)在语言与视觉推理方面表现突出,推动了购物助手、AI软件工程师等目标应用的智能体发展。尽管已有多个数据科学基准被提出,但因设置过于简化,仍无法反映真实应用场景。为此,本文提出DSBench,一个涵盖466个数据分析任务与74个数据建模任务的综合性基准,数据来源为Eloquence和Kaggle竞赛。该基准通过长上下文、多模态背景、大文件与多表结构推理,以及端到端建模任务,实现真实环境模拟。对主流LLMs、LVLMs及智能体的评估显示,它们普遍表现不佳:最优智能体仅解决34.12%的数据分析任务,相对性能差距(RPG)达34.74%。结果表明,开发更实用、智能、自主的数据科学智能体仍需重大突破。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。