测试大模型在真实数据科学任务中的编程能力,发现当前最佳模型仅30.5%准确率。
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
- 基于真实复杂数据的任务设计,需规划与上下文理解能力。
- 使用真实多源数据,覆盖数据清洗与分析全流程。
- 专为评估智能体式数据科学编程而设,适合研究者与开发者参考。
我们提出DA-Code,一个针对大语言模型在智能体式数据科学任务中代码生成能力的评测基准。该基准包含三个核心特点:首先,任务本身具有挑战性,不同于传统代码生成,要求模型具备高级的代码落地与规划能力;其次,所有示例基于真实且多样化的数据,涵盖广泛复杂的数据处理与分析任务;第三,解题需使用复杂的数据科学编程语言完成精细数据操作并得出答案。我们构建了一个可控制、可执行的环境,贴近真实数据分析场景且具备可扩展性。标注者精心设计评估流程以确保评价的准确性与鲁棒性。我们提出了DA-Agent基线模型,实验显示尽管其优于现有框架,但当前最优大模型在该基准上仅达30.5%准确率,表明仍有巨大提升空间。基准已公开于https://da-code-bench.github.io。
原文摘要 · Abstract (English)
We introduce DA-Code, a code generation benchmark specifically designed to assess LLMs on agent-based data science tasks. This benchmark features three core elements: First, the tasks within DA-Code are inherently challenging, setting them apart from traditional code generation tasks and demanding advanced coding skills in grounding and planning. Second, examples in DA-Code are all based on real and diverse data, covering a wide range of complex data wrangling and analytics tasks. Third, to solve the tasks, the models must utilize complex data science programming languages, to perform intricate data processing and derive the answers. We set up the benchmark in a controllable and executable environment that aligns with real-world data analysis scenarios and is scalable. The annotators meticulously design the evaluation suite to ensure the accuracy and robustness of the evaluation. We develop the DA-Agent baseline. Experiments show that although the baseline performs better than other existing frameworks, using the current best LLMs achieves only 30.5% accuracy, leaving ample room for improvement. We release our benchmark at https://da-code-bench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。