arXiv:2608.18740cs.AI2026-08

多智能体系统自动完成企业数据分析与洞察生成,准确率超95%。

A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation

论文配图:A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation
图 1 · 摘自论文原文
  • 五个专业智能体分步处理自然语言查询并生成可视化。
  • 端到端测试准确率达95.3%,响应延迟仅24秒,幻觉率低于7%。
  • 适合需要自动化商业分析的企业和数据平台开发者。

本文提出一个基于CrewAI的多智能体框架,用于对话式商业智能。五个专用智能体按顺序处理自然语言查询,检索与分析数据,通过模型上下文协议(MCP)生成可视化,并输出可操作的洞察。平台采用纵深防御安全架构实现多租户数据隔离,并引入查询参数化机制,将对话洞察转化为可复用的仪表板组件。在300个涵盖合成与生产数据集的端到端测试中,功能准确率达到95.3%,平均响应延迟为24秒,由大模型评估的响应质量得分为4.52/5.0,幻觉率为93.0%,相比单智能体基线提升22.6个百分点准确率与20.2%质量。跨四类大模型评估及人工专家验证确认了架构泛化性与评估可靠性。消融实验证明数据解析与报告聚合智能体是输出质量的核心驱动力。

原文摘要 · Abstract (English)

This paper proposes a multi-agent framework built on CrewAI [1] for conversational business intelligence. Five specialized AI agents operate in a sequential pipeline to process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol (MCP) [2], and deliver actionable insights. The platform features a defense-in-depth security architecture for multi-tenant data isolation and a query parameterization mechanism for transforming conversational insights into reusable dashboard components. Evaluation across 300 end-to-end test cases spanning synthetic and production enterprise datasets demonstrates 95.3% functional accuracy, a mean response latency of 24 seconds, and a response quality score of 4.52/5.0 as assessed by an LLM-as-a-Judge framework, with a 93.0% hallucination-free rate, representing a 22.6 percentage point accuracy improvement and 20.2% quality gain over a single-agent baseline. Cross-model evaluation across four LLM backends and human expert validation confirm architectural generalizability and evaluator reliability. An ablation study confirms that the Data Analysis and Report Aggregation agents are the primary drivers of output quality.

企业分析多智能体对话式BI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。