从传统RAG到智能体RAG,提升企业数据整合的可信与效率
Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG

- 用知识图谱增强RAG,让推理有据可依
- 多智能体协作实现复杂任务的自主规划与验证
- 解决大模型在企业中的成本与准确率难题
大型语言模型(LLMs)和人工智能代理在零样本与少样本场景下展现出强大的数据集成潜力。然而,在企业环境中,由于持续存在的知识鸿沟,其准确性与成本仍面临严峻挑战。本文提出通过基于知识的LLM与代理,在检索增强生成(RAG)工作流中实现可信、可扩展且成本可控的数据集成。其中,可信性指决策基于可验证证据,透明支持、抗幻觉、任务间一致。论文梳理了从经典RAG到GraphRAG与KG-RAG(基于知识图谱的RAG)的发展脉络,揭示其如何融合参数化与上下文知识。在此基础上,探索向智能体式RAG的演进:自主多智能体系统可自适应规划、检索、精炼与推理,应对复杂集成任务。同时分析优化策略以降低大规模企业场景下的计算开销。最后,指出构建可靠、可解释、可扩展的知识驱动集成系统的开放挑战与未来方向。
原文摘要 · Abstract (English)
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG) workflow. Here, trustworthiness refers to evidence-grounded, verifiable reasoning, where integration decisions are transparently supported by retrieved knowledge, robust against hallucination, and consistent across tasks. We trace the evolution from classic RAG to GraphRAG and KG-RAG (knowledge graph-based RAG), highlighting how these paradigms bridge parametric and contextual knowledge. Building on this trajectory, we explore the shift toward Agentic RAG, where autonomous multi-agent systems adaptively plan, retrieve, refine, and reason for complex integration tasks. We examine optimization strategies for cost-efficient integration, addressing computational bottlenecks in large-scale enterprise settings. Finally, we outline open challenges and future directions toward building reliable, explainable, and scalable knowledge-grounded integration systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。