评测4大AI编程助手在数据科学任务中的表现,发现表现参差不齐。
LLM4DS: Evaluating Large Language Models for Data Science Code Generation
- 用真实数据科学任务测试四大AI编程工具,按难度和类型评估
- 仅ChatGPT和Claude超过60%成功率,均未达70%上限
- 适合需要稳定输出的开发者,但对复杂任务仍有限制
大型语言模型(LLM)在数据科学代码生成中具有巨大潜力,可提升数据处理、统计分析与可视化效率。然而其在该领域的实际效果尚不明确。本文通过控制实验,评估了四种主流AI助手——Microsoft Copilot(GPT-4 Turbo)、ChatGPT(o1-preview)、Claude(3.5 Sonnet)和Perplexity Labs(Llama-3.1-70b-instruct)——在来自Stratascratch平台的多样化数据科学编程挑战中的表现。采用目标-问题-度量(GQM)方法,从任务类型(分析型、算法型、可视化型)和难度层级进行评估。结果显示,所有模型均超过50%基准成功率,证明其能力远超随机猜测。其中,仅ChatGPT与Claude的准确率显著高于60%,但无一模型达到70%阈值,表明其在高要求任务上仍有局限。ChatGPT在不同难度下表现稳定,而Claude的成功率随任务复杂度波动。假设检验显示任务类型对成功率无显著影响。针对分析类任务的效率分析表明,执行时间差异不显著,尽管ChatGPT虽成功率高,但耗时更长且结果更不可预测。本研究提供了数据科学领域中对LLM的结构化实证评估,为模型选型提供依据,并强调超越基础准确率的严谨评估的重要性。
原文摘要 · Abstract (English)
The adoption of Large Language Models (LLMs) for code generation in data science offers substantial potential for enhancing tasks such as data manipulation, statistical analysis, and visualization. However, the effectiveness of these models in the data science domain remains underexplored. This paper presents a controlled experiment that empirically assesses the performance of four leading LLM-based AI assistants-Microsoft Copilot (GPT-4 Turbo), ChatGPT (o1-preview), Claude (3.5 Sonnet), and Perplexity Labs (Llama-3.1-70b-instruct)-on a diverse set of data science coding challenges sourced from the Stratacratch platform. Using the Goal-Question-Metric (GQM) approach, we evaluated each model's effectiveness across task types (Analytical, Algorithm, Visualization) and varying difficulty levels. Our findings reveal that all models exceeded a 50% baseline success rate, confirming their capability beyond random chance. Notably, only ChatGPT and Claude achieved success rates significantly above a 60% baseline, though none of the models reached a 70% threshold, indicating limitations in higher standards. ChatGPT demonstrated consistent performance across varying difficulty levels, while Claude's success rate fluctuated with task complexity. Hypothesis testing indicates that task type does not significantly impact success rate overall. For analytical tasks, efficiency analysis shows no significant differences in execution times, though ChatGPT tended to be slower and less predictable despite high success rates. This study provides a structured, empirical evaluation of LLMs in data science, delivering insights that support informed model selection tailored to specific task demands. Our findings establish a framework for future AI assessments, emphasizing the value of rigorous evaluation beyond basic accuracy measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。