arXiv:2608.08445cs.AI2026-08

RAG并非新概念,而是早期信息检索技术的延续与升级。

Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective

论文配图:Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
图 1 · 摘自论文原文
  • 从信息检索历史追溯RAG核心思想的起源
  • 发现2000年代已有类似知识增强与查询优化方法
  • 适合关注AI发展脉络与跨领域融合的研究者

检索增强生成(RAG)常被视为大语言模型(LLM)局限性的产物,是一种将生成结果锚定于外部知识的新范式。然而,从更广阔的历史视角看,这一观点并不完整。本文指出,RAG的核心理念——如检索与生成融合、知识增强、答案验证及迭代查询优化——早在2000年代初的信息检索(IR)与问答(QA)研究中已被探索并实现,远早于大语言模型的出现。我们通过系统追溯现代RAG及其代理型变体的古典IR与QA源头,揭示其长期被忽视的原因:学术社区割裂、术语变迁以及快速演进领域中的“近因偏见”。与其将LLM视为检索增强智能的起点,不如将其视为数十年来问答架构上的新接口层。这种重新定位不仅具有历史意义,更使被遗忘的前期工作——如用户建模、答案验证与查询优化——得以重见天日,直接指导下一代RAG设计,避免重复发明,并推动真正意义上的跨领域整合。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a broader historical context. In this paper, we argue that the core ideas underlying RAG are not new: foundational concepts such as integrating retrieval and language generation, knowledge augmentation, answer verification, and iterative query (or prompt) refinement had already been studied and instantiated in information retrieval (IR) and question answering (QA) research dating back to the early 2000s, well before the emergence of LLMs. We make this case by systematically tracing the intellectual lineage of modern RAG and Agentic RAG back to their classical IR and QA antecedents, and examining why this continuity has gone under-recognized -- a consequence of community fragmentation, shifting terminology, and the recency bias endemic to fast-moving fields. Rather than treating LLMs as the origin point of retrieval-augmented intelligence, we propose viewing them as a new interface layer atop a decades-old QA architecture. This reframing is not merely historical: by situating RAG within the longer trajectory of IR research, we surface underutilized prior work -- on user modeling, answer validation, and query refinement -- that can directly inform next-generation RAG design, reducing unintentional rediscovery and fostering genuine cross-community integration.

RAG信息检索历史脉络知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。