新基准DSAEval评估数据科学智能体在真实任务中的表现。
DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems

- 构建641个真实数据科学问题,覆盖多模态数据与多轮交互。
- Claude-Sonnet-4.5综合表现最佳,多模态感知提升视觉任务准确率2.04%~11.30%。
- 适合研究数据智能体评估、多模态推理与自动化分析的开发者参考。
基于大模型的数据智能体旨在自动化从数据分析到深度学习的任务。然而,真实世界数据科学问题开放性强、跨多个分类体系且无标准答案,给评估带来挑战。为此,我们提出DSAEval基准,包含641个源于285个多样化数据集的真实问题,涵盖结构化与非结构化数据(如图像和文本)。该基准具有三大特性:(1) 多模态环境感知,支持文本与视觉信息理解;(2) 多轮交互设计,模拟真实项目中迭代累积的流程;(3) 多维评估,全面衡量推理、代码与结果质量。我们系统评估了13个先进智能体模型。结果显示,Claude-Sonnet-4.5整体性能最强,MiMo-V2-Pro在耗时效率上领先,GPT-5.2在步骤效率最优,MiMo-V2-Flash成本最低。进一步发现,多模态感知在视觉相关任务中显著提升性能,增益达2.04%至11.30%。总体而言,当前智能体在结构化数据与常规分析流程中表现良好,但在非结构化领域仍面临巨大挑战。最后,我们提供关键洞察并展望未来研究方向。
原文摘要 · Abstract (English)
Recent LLM-based data agents aim to automate data science tasks ranging from data analysis to deep learning. However, the open-ended nature of real-world data science problems, which often span multiple taxonomies and lack standard answers, poses a significant challenge for evaluation. To address this, we introduce DSAEval, a benchmark comprising 641 real-world data science problems grounded in 285 diverse datasets, covering both structured and unstructured data (e.g., image and text). DSAEval incorporates three distinctive features: (1) Multimodal Environment Perception, which enables agents to interpret observations from multiple modalities, including text and vision; (2) Multi-Query Interactions, which mirror the iterative and cumulative nature of real-world data science projects; and (3) Multi-Dimensional Evaluation, which provides a holistic assessment across reasoning, code, and results. We systematically evaluate 13 recent advanced agentic LLMs using DSAEval. Our results show that Claude-Sonnet-4.5 achieves the strongest overall performance, MiMo-V2-Pro and GPT-5.2 lead in duration and step efficiency, respectively, and MiMo-V2-Flash is the most cost-effective. We further demonstrate that multimodal perception consistently improves performance on vision-related tasks, with gains ranging from 2.04\% to 11.30\%. Overall, while current data science agents perform well on structured data and routine data analysis workflows, substantial challenges remain in unstructured domains. Finally, we offer critical insights and outline future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。