综述大模型驱动的数据科学智能体能力与挑战,梳理全流程自动化路径。
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
- 构建六阶段数据科学流程的分类体系,覆盖从理解到部署全周期。
- 45个系统中90%缺乏安全机制,多模态推理与工具调度仍存瓶颈。
- 适合关注AI自动化、可信智能体研发的研究者与工程师参考。
大语言模型(LLMs)的进步催生了一类新型AI智能体,能够通过整合规划、工具使用和跨文本、代码、表格、视觉等多模态推理,自动完成数据科学工作流的多个阶段。本综述首次提出面向生命周期的数据科学智能体全面分类框架,系统分析并映射了45个现有系统至端到端数据科学流程的六个阶段:业务理解与数据获取、探索性分析与可视化、特征工程、模型构建与选择、解释与说明、部署与监控。除流程覆盖外,还从五个跨领域设计维度对每个智能体进行标注:推理与规划方式、模态融合能力、工具编排深度、学习与对齐方法,以及信任、安全与治理机制。在分类基础上,我们批判性总结智能体能力,指出各阶段的优势与局限,并回顾新兴基准与评估实践。分析揭示三大趋势:多数系统聚焦探索分析、可视化与建模,而忽视业务理解、部署与监控;多模态推理与工具编排仍是未解难题;超过90%的系统缺乏显式信任与安全机制。最后,我们提出对齐稳定性、可解释性、治理与鲁棒评估框架等开放挑战,并为未来研究提供方向,推动开发更稳健、可信、低延迟、透明且广泛可用的数据科学智能体。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled a new class of AI agents that automate multiple stages of the data science workflow by integrating planning, tool use, and multimodal reasoning across text, code, tables, and visuals. This survey presents the first comprehensive, lifecycle-aligned taxonomy of data science agents, systematically analyzing and mapping forty-five systems onto the six stages of the end-to-end data science process: business understanding and data acquisition, exploratory analysis and visualization, feature engineering, model building and selection, interpretation and explanation, and deployment and monitoring. In addition to lifecycle coverage, we annotate each agent along five cross-cutting design dimensions: reasoning and planning style, modality integration, tool orchestration depth, learning and alignment methods, and trust, safety, and governance mechanisms. Beyond classification, we provide a critical synthesis of agent capabilities, highlight strengths and limitations at each stage, and review emerging benchmarks and evaluation practices. Our analysis identifies three key trends: most systems emphasize exploratory analysis, visualization, and modeling while neglecting business understanding, deployment, and monitoring; multimodal reasoning and tool orchestration remain unresolved challenges; and over 90% lack explicit trust and safety mechanisms. We conclude by outlining open challenges in alignment stability, explainability, governance, and robust evaluation frameworks, and propose future research directions to guide the development of robust, trustworthy, low-latency, transparent, and broadly accessible data science agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。