arXiv:2506.03100cs.LGcs.AI2025-06被引 4

首次为RAG提供理论分析,揭示其与上下文学习的内在联系及泛化误差上限。

Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds

  • 将检索内容视为带噪声的上下文样本,统一建模RAG与ICL
  • 推导出RAG在有限样本下的泛化误差界,发现其存在固有误差天花板
  • 适用于关注大模型可解释性与知识增强机制的研究者

近年来,检索增强生成(RAG)通过引入外部知识在大语言模型中取得了诸多实证成功,但其理论研究仍不充分。本文首次为基于上下文线性回归的RAG推导出有限样本泛化界,并揭示了精确的偏差-方差权衡关系。我们的框架将检索到的文本视为依赖查询的噪声上下文样例,恢复了经典上下文学习(ICL)和标准RAG作为极限情形。分析表明,相较于ICL,RAG存在固有的泛化误差上限。此外,该框架可通过引入均匀与非均匀的RAG噪声,同时建模从训练数据和外部语料库中检索的情形。实验验证了理论预测,展示了ICL与RAG在Natural Questions和TriviaQA等常见问答基准上的样本效率。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) has seen many empirical successes in recent years by aiding the LLM with external knowledge. However, its theoretical aspect has remained mostly unexplored. In this paper, we propose the first finite-sample generalization bound for RAG in in-context linear regression and derive an exact bias-variance tradeoff. Our framework views the retrieved texts as query-dependent noisy in-context examples and recovers the classical in-context learning (ICL) and standard RAG as the limit cases. Our analysis suggests that an intrinsic ceiling on generalization error exists on RAG as opposed to the ICL. Furthermore, our framework is able to model retrieval both from the training data and from external corpora by introducing uniform and non-uniform RAG noise. In line with our theory, we show the sample efficiency of ICL and RAG empirically with experiments on common QA benchmarks, such as Natural Questions and TriviaQA.

RAG理论分析上下文学习泛化误差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。