综述大模型如何赋能数据科学全流程自动化
Large Language Model-based Data Science Agent: A Survey
- 从代理设计与数据科学流程双视角构建分析框架
- 系统梳理大模型在数据清洗、建模、评估等环节的应用
- 适合关注AI+数据科学融合的从业者和研究者
大型语言模型(LLMs)的快速发展推动了多个领域的创新应用,基于大模型的智能体正成为重要研究方向。本文全面综述了面向数据科学任务的LLM智能体,总结了近期研究成果。从智能体视角出发,探讨了角色设定、执行机制、知识管理与反思方法等关键设计原则;从数据科学视角出发,识别出数据预处理、模型开发、评估与可视化等核心流程。本文贡献包括:(1) 系统性回顾了大模型智能体在数据科学任务中的最新进展;(2) 提出一个连接通用智能体设计与数据科学实际工作流的双重视角框架。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has driven novel applications across diverse domains, with LLM-based agents emerging as a crucial area of exploration. This survey presents a comprehensive analysis of LLM-based agents designed for data science tasks, summarizing insights from recent studies. From the agent perspective, we discuss the key design principles, covering agent roles, execution, knowledge, and reflection methods. From the data science perspective, we identify key processes for LLM-based agents, including data preprocessing, model development, evaluation, visualization, etc. Our work offers two key contributions: (1) a comprehensive review of recent developments in applying LLMbased agents to data science tasks; (2) a dual-perspective framework that connects general agent design principles with the practical workflows in data science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。