arXiv:2509.23988cs.AIcs.DB2025-09综述被引 31

大模型和智能体让数据分析师实现自动化,支持多模态数据分析。

LLM/Agent-as-Data-Analyst: A Survey

  • 用大模型+智能体实现自然语言驱动的数据分析
  • 覆盖结构化、半结构化、非结构化及异构数据全场景
  • 适合需要自动分析数据的科研与企业用户

大型语言模型(LLMs)与智能体技术正重塑数据分析的功能与开发范式(即 LLM/Agent-as-Data-Analyst),在学术界与工业界均产生深远影响。相比传统规则或小模型方法,(智能体)大模型可实现复杂数据理解、自然语言接口、语义分析功能及自主流程编排。从模态角度,本文综述了基于大模型的技术在(i)结构化数据(如 NL2SQL、NL2GQL、ModelQA)、(ii)半结构化数据(如标记语言理解、半结构化表格问答)、(iii)非结构化数据(如图表理解、文本/图像文档理解)以及(iv)异构数据(如数据湖中的数据检索与模态对齐)中的应用。技术演进进一步提炼出智能数据分析代理的四大设计目标:语义感知、自主流程、工具增强工作流、开放世界任务支持。最后,本文指出现存挑战,并提出若干推进方向。

原文摘要 · Abstract (English)

Large language models (LLMs) and agent techniques have brought a fundamental shift in the functionality and development paradigm of data analysis tasks (a.k.a LLM/Agent-as-Data-Analyst), demonstrating substantial impact across both academia and industry. In comparison with traditional rule or small-model based approaches, (agentic) LLMs enable complex data understanding, natural language interfaces, semantic analysis functions, and autonomous pipeline orchestration. From a modality perspective, we review LLM-based techniques for (i) structured data (e.g., NL2SQL, NL2GQL, ModelQA), (ii) semi-structured data (e.g., markup languages understanding, semi-structured table question answering), (iii) unstructured data (e.g., chart understanding, text/image document understanding), and (iv) heterogeneous data (e.g., data retrieval and modality alignment in data lakes). The technical evolution further distills four key design goals for intelligent data analysis agents, namely semantic-aware design, autonomous pipelines, tool-augmented workflows, and support for open-world tasks. Finally, we outline the remaining challenges and propose several insights and practical directions for advancing LLM/Agent-powered data analysis.

大模型数据分析智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。