arXiv:2602.24288cs.AIcs.CL2026-02被引 2

DARE-bench评估大模型在数据科学任务中的指令遵循与流程准确性,填补了评测标准空白。

DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science

  • 基于可验证真值构建6300个真实数据科学任务,实现客观评估。
  • 微调后模型准确率最高提升8.3倍,证明训练数据关键性。
  • 适合研究模型泛化能力、训练数据设计及自动化数据分析工具开发者。

大型语言模型在处理复杂多步骤数据科学任务时需求激增,现有基准存在两大缺陷:缺乏标准化的过程感知评估以捕捉指令遵循与流程保真度,以及高质量标注训练数据稀缺。为此,我们提出DARE-bench,一个面向机器学习建模与数据科学指令遵循的基准。不同于依赖人工或模型评分的基准,DARE-bench所有任务均具备可验证真值,确保评估客观可复现。该基准包含6300个源自Kaggle的任务,涵盖广泛场景并支持智能体工具使用,提供大规模训练与评估数据集。大量实验表明,即使性能较强的gpt-o4-mini在建模任务上表现不佳。利用DARE-bench训练数据进行微调可显著提升模型性能:监督微调使Qwen3-32B准确率提升1.83倍,强化学习使Qwen3-4B准确率提升超过8倍。这些结果验证了DARE-bench作为精准评估工具与核心训练数据的重要价值。

原文摘要 · Abstract (English)

The fast-growing demands in using Large Language Models (LLMs) to tackle complex multi-step data science tasks create an emergent need for accurate benchmarking. There are two major gaps in existing benchmarks: (i) the lack of standardized, process-aware evaluation that captures instruction adherence and process fidelity, and (ii) the scarcity of accurately labeled training data. To bridge these gaps, we introduce DARE-bench, a benchmark designed for machine learning modeling and data science instruction following. Unlike many existing benchmarks that rely on human- or model-based judges, all tasks in DARE-bench have verifiable ground truth, ensuring objective and reproducible evaluation. To cover a broad range of tasks and support agentic tools, DARE-bench consists of 6,300 Kaggle-derived tasks and provides both large-scale training data and evaluation sets. Extensive evaluations show that even highly capable models such as gpt-o4-mini struggle to achieve good performance, especially in machine learning modeling tasks. Using DARE-bench training tasks for fine-tuning can substantially improve model performance. For example, supervised fine-tuning boosts Qwen3-32B's accuracy by 1.83x and reinforcement learning boosts Qwen3-4B's accuracy by more than 8x. These significant improvements verify the importance of DARE-bench both as an accurate evaluation benchmark and critical training data.

大模型评测数据科学指令遵循训练数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。