揭秘RAG系统安全隐私风险,构建可信生成体系
Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems

- 划分检索、上下文构建、生成三阶段威胁面,系统梳理攻击类型
- 揭示查询日志、索引、知识库污染等多路径信息泄露风险
- 适合关注AI安全、可信生成的开发者与研究者参考
检索增强生成(RAG)通过结合检索机制与生成模型,显著提升大语言模型的事实准确性与跨领域适应性。然而,引入检索流程也带来了新的安全与隐私挑战,敏感信息可能通过检索索引、查询日志、上下文构造或联邦更新等环节泄露,而知识库的对抗性操纵更会破坏生成内容的可信度。本文全面分析了集中式、本地部署(Micro-RAG)、联邦及混合模式下RAG系统的隐私与安全风险,提出涵盖检索、上下文构建与生成阶段的统一威胁分类框架,系统梳理了成员推断、索引推断、投毒、梯度泄露与合谋攻击等攻击类型。同时综述了架构、算法与密码学层面的防御策略,强调隐私与性能的权衡及部署考量。最后指出未来构建可信、安全、鲁棒RAG系统的关键开放问题。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) has emerged as a dominant paradigm for enhancing large language models with external knowledge. By coupling retrieval mechanisms with generative models, RAG systems improve factual grounding and adaptability across domains. However, integrating retrieval pipelines introduces new security and privacy risks that extend beyond conventional language modeling threats. Sensitive information may be exposed through retrieval indices, query logs, context construction, or federated updates, while adversarial manipulation of knowledge bases can undermine trust in generated outputs. This survey provides a comprehensive examination of privacy and security challenges across RAG systems deployed in centralized, on-device (Micro-RAG), federated, and hybrid paradigms. We present a unified taxonomy of threat surfaces spanning the retrieval, context construction, and generation stages and systematically analyze attack classes, including membership inference, index inference, poisoning, gradient leakage, and collusion. We further review architectural, algorithmic, and cryptographic defenses, highlighting privacy-utility trade-offs and deployment considerations. Finally, we outline open research challenges toward building trustworthy, secure, and resilient RAG systems for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。