用户用自然语言提问,系统自动跨领域分析结构化与非结构化数据。
AgenticData: An Agentic Data Analytics System for Heterogeneous Data
- 通过多智能体协作生成并优化语义分析计划
- 在三个基准上准确率显著超越现有方法
- 适合非专家快速完成复杂数据分析任务
现有非结构化数据解析系统依赖专家编写代码并管理复杂分析流程,成本高且耗时。为此,我们提出AgenticData,一种创新的智能体式数据分析系统,用户只需提出自然语言问题,系统即可自主分析跨多个领域的数据源,涵盖结构化与非结构化数据。首先,AgenticData采用反馈驱动的规划技术,将自然语言查询自动转化为由关系与语义操作符组成的语义计划。我们设计多智能体协作策略:数据探查智能体用于发现相关数据,语义交叉验证智能体基于反馈迭代优化,智能记忆体则维持短期上下文与长期知识。其次,我们提出语义优化模型,有效精炼并执行语义计划。系统在三个基准上进行了测试,实验结果表明,AgenticData在简单与复杂任务上均取得更优准确率,显著优于现有最先进方法。
原文摘要 · Abstract (English)
Existing unstructured data analytics systems rely on experts to write code and manage complex analysis workflows, making them both expensive and time-consuming. To address these challenges, we introduce AgenticData, an innovative agentic data analytics system that allows users to simply pose natural language (NL) questions while autonomously analyzing data sources across multiple domains, including both unstructured and structured data. First, AgenticData employs a feedback-driven planning technique that automatically converts an NL query into a semantic plan composed of relational and semantic operators. We propose a multi-agent collaboration strategy by utilizing a data profiling agent for discovering relevant data, a semantic cross-validation agent for iterative optimization based on feedback, and a smart memory agent for maintaining short-term context and long-term knowledge. Second, we propose a semantic optimization model to refine and execute semantic plans effectively. Our system, AgenticData, has been tested using three benchmarks. Experimental results showed that AgenticData achieved superior accuracy on both easy and difficult tasks, significantly outperforming state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。