arXiv:2608.24977cs.CRcs.CL2026-08中稿 · Findings of the As…综述

系统梳理RAG的攻击与防御,揭示知识增强模型的安全隐患

Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation

论文配图:Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation
图 1 · 摘自论文原文
  • 按检索、重排、生成各阶段划分攻击与防御策略
  • 发现三大威胁:误导准确性、泄露隐私、破坏公平性
  • 适合关注大模型安全的研究者和开发者

检索增强生成(RAG)通过引入外部知识提升大语言模型的准确性并减少幻觉。然而,这一流程也带来了新的鲁棒性与安全风险,包括语料库投毒、后门攻击、隐私泄露及公平性问题。尽管该领域进展迅速,现有综述仍未能全面覆盖攻击目标、威胁模型及全链路阶段的防御措施。本文提出一个统一且面向流水线的RAG鲁棒性分析框架,对语料库、检索器与生成器三个环节分别建模威胁,并将攻击归纳为准确率、隐私与公平性三类核心目标。同时从流水线视角回顾防御技术,涵盖检索、重排、生成与溯源阶段。此外,总结了鲁棒性评估基准与可解释性方法,以更深入地评测和理解RAG系统的安全性。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. At the same time, the retrieval-augmented pipeline introduces new robustness and security risks, including corpus poisoning, backdoor attacks, privacy leakage, and fairness violations. Despite rapid progress in this area, existing surveys remain limited in their treatment of attacker objectives, threat models, and stage-specific defenses across the full RAG pipeline. This survey presents a unified and pipeline-aware overview of RAG robustness. We formalize threat models over the corpus, retriever, and generator, and organize attacks into three main objectives: accuracy, privacy, and fairness. We further review defenses from a pipeline-aware perspective, covering the retrieval, rerank, generation, and traceback stages. In addition, we summarize robustness benchmarks and explainability methods for more deeply evaluating and explaining RAG robustness.

RAG安全攻击防御大模型可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。