arXiv:2412.14222cs.AIcs.CL2024-12综述被引 52

大模型驱动的数据分析代理,让普通人也能轻松处理复杂数据任务。

A Survey on Large Language Model-based Agents for Statistics and Data Science

  • 用大模型构建可自主规划、推理与协作的数据代理框架
  • 支持多场景应用,减少人工干预,提升数据分析效率
  • 适合非专业人士快速上手数据科学,也启发未来智能分析工具研发

近年来,基于大语言模型(LLM)的“数据代理”展现出变革传统数据分析范式的重要潜力。本综述系统梳理了基于LLM的数据代理的发展脉络、核心能力与实际应用,强调其在简化复杂数据任务、降低用户门槛方面的价值。我们深入探讨当前主流框架的设计趋势,涵盖规划、推理、反思、多代理协作、人机交互、知识融合与系统架构等关键特性,这些机制使代理能在极少人工干预下解决数据驱动问题。此外,通过多个案例研究展示数据代理在真实场景中的应用效果。最后,识别当前面临的核心挑战,并提出未来研究方向,推动数据代理向智能化统计分析软件演进。

原文摘要 · Abstract (English)

In recent years, data science agents powered by Large Language Models (LLMs), known as "data agents," have shown significant potential to transform the traditional data analysis paradigm. This survey provides an overview of the evolution, capabilities, and applications of LLM-based data agents, highlighting their role in simplifying complex data tasks and lowering the entry barrier for users without related expertise. We explore current trends in the design of LLM-based frameworks, detailing essential features such as planning, reasoning, reflection, multi-agent collaboration, user interface, knowledge integration, and system design, which enable agents to address data-centric problems with minimal human intervention. Furthermore, we analyze several case studies to demonstrate the practical applications of various data agents in real-world scenarios. Finally, we identify key challenges and propose future research directions to advance the development of data agents into intelligent statistical analysis software.

大模型代理数据科学智能分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。