arXiv:2601.16344cs.AI2026-01被引 12

构建可扩展的自动化数据科学评估框架,解决现有基准缺陷。

DSGym: A Holistic Framework for Evaluating and Training Data Science Agents

  • 提供模块化架构,支持任务与工具动态添加。
  • 构建2000例训练数据,40亿参数模型超越GPT-4o表现。
  • 聚焦真实数据分析流程,避免虚假捷径解法。

数据科学代理有望通过将数据转化为可执行分析和发现来加速科研进程。然而,现有数据科学基准存在评价接口碎片化、任务覆盖窄及缺乏严格数据依赖性等问题。我们发现,当前多数基准任务可在不使用真实数据的情况下完成。为此,本文提出DSGym,一个标准化的、自包含执行环境下的数据科学代理评估与训练框架。不同于静态基准,DSGym采用模块化设计,便于扩展任务、代理模板与工具,具备持续演进能力。我们构建了DSGym-Tasks,通过质量筛选与捷径可解性检测,对现有基准进行重构与优化。进一步拓展覆盖范围:(1) DSBio——基于文献的专家级生物信息学任务;(2) DSPredict——涵盖计算机视觉、分子预测与单细胞扰动等领域的挑战性预测任务。除评估外,DSGym支持通过执行验证的数据合成流水线进行代理训练。以案例研究为例,我们构建2000条训练样本,训练一个40亿参数模型,在标准化分析基准上表现优于GPT-4o。总体而言,DSGym实现了在真实科研场景下对代理规划、实现与验证数据分析能力的端到端严格评估。

原文摘要 · Abstract (English)

Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short due to fragmented evaluation interfaces that make cross-benchmark comparison difficult, narrow task coverage and a lack of rigorous data grounding. In particular, we show that a substantial portion of tasks in current benchmarks can be solved without using the actual data. To address these limitations, we introduce DSGym, a standardized framework for evaluating and training data science agents in self-contained execution environments. Unlike static benchmarks, DSGym provides a modular architecture that makes it easy to add tasks, agent scaffolds, and tools, positioning it as a live, extensible testbed. We curate DSGym-Tasks, a holistic task suite that standardizes and refines existing benchmarks via quality and shortcut solvability filtering. We further expand coverage with (1) DSBio: expert-derived bioinformatics tasks grounded in literature and (2) DSPredict: challenging prediction tasks spanning domains such as computer vision, molecular prediction, and single-cell perturbation. Beyond evaluation, DSGym enables agent training via execution-verified data synthesis pipeline. As a case study, we build a 2,000-example training set and trained a 4B model in DSGym that outperforms GPT-4o on standardized analysis benchmarks. Overall, DSGym enables rigorous end-to-end measurement of whether agents can plan, implement, and validate data analyses in realistic scientific context.

数据科学智能代理评估框架训练数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。