arXiv:2507.23336cs.AIcs.CL2025-07被引 2

构建真实场景下的数据科学任务评估基准,测试大模型实际表现。

DSBC : Data Science task Benchmarking with Context engineering

  • 基于真实用户交互数据设计上下文工程评测框架
  • 三款大模型在八类任务中表现差异显著,温度参数影响结果
  • 适合研究智能数据分析代理的开发者和评估者参考

大型语言模型(LLMs)的进展深刻影响了数据科学工作流程,催生了专用于自动化分析任务的数据科学代理。尽管应用迅速普及,系统性评估这些代理效能与局限性的基准仍十分稀缺。本文提出一个全面的基准,通过观察我们商业应用中的真实用户交互行为来反映实际使用场景。评估了三种LLM:Claude-4.0-Sonnet、Gemini-2.5-Flash 和 OpenAI-o4-Mini,采用三种方法:零样本+上下文工程、多步+上下文工程,以及 SmolAgent。基准涵盖八类数据科学任务,并探究模型对常见提示问题(如数据泄露、指令模糊)的敏感性。同时研究温度参数对各模型及任务层面结果的影响。结果显示不同模型与方法间存在明显性能差异,揭示了实际部署中的关键影响因素。本文提出的基准数据集与评估框架旨在为未来更鲁棒、高效的数据科学代理研究提供基础。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have significantly impacted data science workflows, giving rise to specialized data science agents designed to automate analytical tasks. Despite rapid adoption, systematic benchmarks evaluating the efficacy and limitations of these agents remain scarce. In this paper, we introduce a comprehensive benchmark specifically crafted to reflect real-world user interactions with data science agents by observing usage of our commercial applications. We evaluate three LLMs: Claude-4.0-Sonnet, Gemini-2.5-Flash, and OpenAI-o4-Mini across three approaches: zero-shot with context engineering, multi-step with context engineering, and with SmolAgent. Our benchmark assesses performance across a diverse set of eight data science task categories, additionally exploring the sensitivity of models to common prompting issues, such as data leakage and slightly ambiguous instructions. We further investigate the influence of temperature parameters on overall and task-specific outcomes for each model and approach. Our findings reveal distinct performance disparities among the evaluated models and methodologies, highlighting critical factors that affect practical deployment. The benchmark dataset and evaluation framework introduced herein aim to provide a foundation for future research of more robust and effective data science agents.

大模型评估数据科学代理上下文工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。