arXiv:2506.23719cs.LGcs.AI2025-06被引 43

评测大模型在真实多步数据任务中的分析能力,发现顶尖模型准确率仅14.55%。

DABstep: Data Agent Benchmark for Multi-step Reasoning

  • 基于真实金融平台构建450+多步数据任务,融合代码与文档推理。
  • 最先进模型在最难任务上准确率仅14.55%,暴露显著能力差距。
  • 提供自动评分与公开排行榜,推动自主数据分析研究。

我们提出DABstep,一个针对真实多步数据分析任务的AI智能体评估基准。该基准包含超过450个源自金融分析平台的真实挑战,要求模型结合代码式数据处理与异构文档的上下文推理。每项任务需迭代式多步求解,考察数据操作、多源交叉验证与精准结果输出能力。基准采用事实型答案格式,并支持自动正确性校验,实现规模化客观评分。我们评估了主流大模型智能体,发现即使最优模型在最困难任务上的准确率也仅为14.55%。本文详述基准设计、数据构成、任务形式与评估协议,报告基线结果并分析失败模式。DABstep已开放公共排行榜与工具包,以加速自主数据分析研究。

原文摘要 · Abstract (English)

We introduce DABstep, a novel benchmark for evaluating AI agents on realistic multi-step data analysis tasks. DABstep comprises over 450 real-world challenges derived from a financial analytics platform, requiring models to combine code-based data processing with contextual reasoning over heterogeneous documentation. Each task demands an iterative, multi-step problem-solving approach, testing capabilities in data manipulation, cross-referencing multiple sources, and precise result reporting. The benchmark provides a factoid-style answer format with automatic correctness checks for objective scoring at scale. We evaluate leading LLM-based agents, revealing a substantial performance gap: even the best agent achieves only 14.55% accuracy on the hardest tasks. We detail our benchmark's design, dataset composition, task formulation, evaluation protocol, report baseline results and analyze failure modes. DABstep is released with a public leaderboard and toolkit to accelerate research in autonomous data analysis.

智能体评测多步推理数据科学大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。