梳理数据科学自动化评估工具,揭示现有研究的三大盲区。
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
- 系统调研数据科学中LLM助手与代理的评估方法
- 发现多数研究聚焦少数目标导向任务,忽视数据管理与探索
- 指出评估偏重完全替代人类,忽略人机协作与任务重构的潜力
数据科学旨在从数据中提取洞察以支持决策。近年来,大语言模型(LLMs)被用作数据科学助手,提供想法、技术建议、代码片段,或对结果进行解释与报告。随着具备代码执行和知识库等能力的LLM代理(LLM agents)兴起,部分数据科学活动的自动化成为可能——这类系统能自主行动并与数字环境交互。本文对数据科学领域中LLM助手与代理的评估进行了综述。研究发现:(1)评估集中于少数目标导向任务,严重忽视数据管理和探索性分析;(2)研究多关注纯辅助或完全自治的代理,缺乏对中间层级人机协作的关注;(3)评估普遍强调取代人类,忽略了通过任务重构实现更高层次自动化的可能性。
原文摘要 · Abstract (English)
Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data science, by suggesting ideas, techniques and small code snippets, or for the interpretation of results and reporting. Proper automation of some data-science activities is now promised by the rise of LLM agents, i.e., AI systems powered by an LLM equipped with additional affordances--such as code execution and knowledge bases--that can perform self-directed actions and interact with digital environments. In this paper, we survey the evaluation of LLM assistants and agents for data science. We find (1) a dominant focus on a small subset of goal-oriented activities, largely ignoring data management and exploratory activities; (2) a concentration on pure assistance or fully autonomous agents, without considering intermediate levels of human-AI collaboration; and (3) an emphasis on human substitution, therefore neglecting the possibility of higher levels of automation thanks to task transformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。