首次系统梳理智能体式检索增强生成的架构与风险,为可信自主系统奠基。
SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions
- 将智能体检索生成视为马尔可夫决策过程,形式化其决策循环
- 揭示幻觉传播、记忆污染等四大系统性风险,挑战现有评估方法
- 提出模块化架构分类与稳定可控的未来研究方向,适合系统设计者参考
检索增强生成(RAG)系统正演变为具备自主性的智能体架构,大语言模型能自主协调多步推理、动态记忆管理与迭代检索策略。尽管工业界快速采纳,当前研究缺乏对这类自主系统的系统性理解,导致架构碎片化、评估方法不一致,且存在严重可靠性风险。本文首次提出统一框架,将智能体检索-生成循环形式化为有限时域部分可观测马尔可夫决策过程,显式建模控制策略与状态转移。基于此,构建了涵盖规划机制、检索编排、记忆范式与工具调用行为的完整分类体系。分析指出传统静态评估存在根本缺陷,识别出幻觉累积传播、记忆污染、检索错位及工具执行级联漏洞等关键风险。最后,提出包括稳定自适应检索、成本感知编排、形式化轨迹评估与监督机制在内的博士级研究方向,为构建可靠、可控、可扩展的智能体检索系统提供明确路线图。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems are increasingly evolving into agentic architectures where large language models autonomously coordinate multi-step reasoning, dynamic memory management, and iterative retrieval strategies. Despite rapid industrial adoption, current research lacks a systematic understanding of Agentic RAG as a sequential decision-making system, leading to highly fragmented architectures, inconsistent evaluation methodologies, and unresolved reliability risks. This Systematization of Knowledge (SoK) paper provides the first unified framework for understanding these autonomous systems. We formalize agentic retrieval-generation loops as finite-horizon partially observable Markov decision processes, explicitly modeling their control policies and state transitions. Building upon this formalization, we develop a comprehensive taxonomy and modular architectural decomposition that categorizes systems by their planning mechanisms, retrieval orchestration, memory paradigms, and tool-invocation behaviors. We further analyze the critical limitations of traditional static evaluation practices and identify severe systemic risks inherent to autonomous loops, including compounding hallucination propagation, memory poisoning, retrieval misalignment, and cascading tool-execution vulnerabilities. Finally, we outline key doctoral-scale research directions spanning stable adaptive retrieval, cost-aware orchestration, formal trajectory evaluation, and oversight mechanisms, providing a definitive roadmap for building reliable, controllable, and scalable agentic retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。