系统梳理RAG安全威胁与防御,构建首个端到端安全评估框架。
Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks
- 按检索生成流程分析数据投毒、对抗攻击等核心威胁
- 提出输入输出双阶段防御体系,涵盖加密检索与差分隐私等技术
- 首次整合权威数据集与评测标准,助力安全实验设计
检索增强生成(RAG)通过引入外部知识库显著缓解大模型的幻觉和领域知识不足问题。然而,其多模块架构带来了复杂的系统级安全漏洞。本文基于RAG工作流,分析潜在漏洞机制,系统分类数据投毒、对抗攻击、成员推断攻击等核心威胁。在此基础上,从输入与输出双重角度构建防御技术分类体系:输入侧涵盖动态访问控制、同态加密检索、对抗预过滤;输出侧总结联邦学习隔离、差分隐私扰动、轻量级数据清洗等防泄漏技术。为统一未来实验设计,整合权威测试数据集、安全标准与评估框架。据我们所知,这是首个专注于RAG系统安全的完整综述。不同于现有孤立研究特定漏洞的文献,本文全面映射整个管道,提供威胁模型、防御机制与评测基准的一体化分析,旨在揭示潜在风险,推动下一代高鲁棒、可信RAG系统的研发。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) significantly mitigates the hallucinations and domain knowledge deficiency in large language models by incorporating external knowledge bases. However, the multi-module architecture of RAG introduces complex system-level security vulnerabilities. Guided by the RAG workflow, this paper analyzes the underlying vulnerability mechanisms and systematically categorizes core threat vectors such as data poisoning, adversarial attacks, and membership inference attacks. Based on this threat assessment, we construct a taxonomy of RAG defense technologies from a dual perspective encompassing both input and output stages. The input-side analysis reviews data protection mechanisms including dynamic access control, homomorphic encryption retrieval, and adversarial pre-filtering. The output-side examination summarizes advanced leakage prevention techniques such as federated learning isolation, differential privacy perturbation, and lightweight data sanitization. To establish a unified benchmark for future experimental design, we consolidate authoritative test datasets, security standards, and evaluation frameworks. To the best of our knowledge, this paper presents the first end-to-end survey dedicated to the security of RAG systems. Distinct from existing literature that isolates specific vulnerabilities, we systematically map the entire pipeline-providing a unified analysis of threat models, defense mechanisms, and evaluation benchmarks. By enabling deep insights into potential risks, this work seeks to foster the development of highly robust and trustworthy next-generation RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。