系统梳理RAG发展脉络,揭示其提升问答准确性的关键技术与挑战
A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
- 整合检索与生成模型,通过外部知识缓解大模型幻觉
- 对比评估多种RAG系统,发现检索精度与生成流畅性存在权衡
- 适合关注大模型可靠性与企业落地的NLP研究者参考
检索增强生成(RAG)是自然语言处理的重要进展,通过结合大语言模型(LLMs)与信息检索系统,提升事实准确性与上下文相关性。本文对RAG进行系统综述,从开放域问答早期发展到当前多样化应用的最新实现,梳理其演进路径。重点分析检索机制、序列生成模型及融合策略等核心组件。按年份梳理关键里程碑与研究趋势,展现RAG的快速演进。探讨企业级部署中的挑战,包括专有数据检索、安全性和可扩展性问题。通过基准测试对比不同RAG实现,评估其在检索精度、生成流畅性、延迟和计算效率方面的表现。持续存在的挑战包括检索质量、隐私顾虑与集成开销。最后提出新兴解决方案,如混合检索、隐私保护技术、优化融合策略及代理式RAG架构,预示更可靠、高效、上下文感知的知识密集型NLP系统未来。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) represents a major advancement in natural language processing (NLP), combining large language models (LLMs) with information retrieval systems to enhance factual grounding, accuracy, and contextual relevance. This paper presents a comprehensive systematic review of RAG, tracing its evolution from early developments in open domain question answering to recent state-of-the-art implementations across diverse applications. The review begins by outlining the motivations behind RAG, particularly its ability to mitigate hallucinations and outdated knowledge in parametric models. Core technical components-retrieval mechanisms, sequence-to-sequence generation models, and fusion strategies are examined in detail. A year-by-year analysis highlights key milestones and research trends, providing insight into RAG's rapid growth. The paper further explores the deployment of RAG in enterprise systems, addressing practical challenges related to retrieval of proprietary data, security, and scalability. A comparative evaluation of RAG implementations is conducted, benchmarking performance on retrieval accuracy, generation fluency, latency, and computational efficiency. Persistent challenges such as retrieval quality, privacy concerns, and integration overhead are critically assessed. Finally, the review highlights emerging solutions, including hybrid retrieval approaches, privacy-preserving techniques, optimized fusion strategies, and agentic RAG architectures. These innovations point toward a future of more reliable, efficient, and context-aware knowledge-intensive NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。