arXiv:2601.03979cs.CRcs.CL2026-01中稿 · publication at the…被引 9

系统梳理RAG系统隐私风险与防护方法,为安全应用提供指南。

SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems

  • 构建RAG隐私风险分类体系与防护流程图
  • 归纳多种隐私攻击类型及对应缓解技术
  • 适合关注AI安全与数据合规的研究者与开发者

大型语言模型(LLMs)在自然语言理解与生成方面展现出巨大潜力,推动了其各类应用场景的快速发展。为弥补LLM固有知识的局限性,检索增强生成(RAG)技术广泛流行。RAG通过将LLM与领域知识库结合,使回答用户问题时能引入上下文和最新信息。然而,随着RAG在敏感数据上的应用增加,隐私泄露风险日益凸显。近期多项研究探讨了RAG中的隐私威胁,包括对抗攻击及缓解策略。本文通过系统性文献综述,首次对RAG隐私风险、缓解技术与评估方法进行系统化整理,提出一个完整的隐私风险分类体系与防护流程图。研究揭示了当前缓解措施的成熟度,并指出了关键注意事项,为提升RAG系统的隐私安全性提供了理论基础与实践参考。

原文摘要 · Abstract (English)

The continued promise of Large Language Models (LLMs), particularly in their natural language understanding and generation capabilities, has driven a rapidly increasing interest in identifying and developing LLM use cases. In an effort to complement the ingrained "knowledge" of LLMs, Retrieval-Augmented Generation (RAG) techniques have become widely popular. At its core, RAG involves the coupling of LLMs with domain-specific knowledge bases, whereby the generation of a response to a user question is augmented with contextual and up-to-date information. The proliferation of RAG has sparked concerns about data privacy, particularly with the inherent risks that arise when leveraging databases with potentially sensitive information. Numerous recent works have explored various aspects of privacy risks in RAG systems, from adversarial attacks to proposed mitigations. With the goal of surveying and unifying these works, we ask one simple question: What are the privacy risks in RAG, and how can they be measured and mitigated? To answer this question, we conduct a systematic literature review of RAG works addressing privacy, and we systematize our findings into a comprehensive set of privacy risks, mitigation techniques, and evaluation strategies. We supplement these findings with two primary artifacts: a Taxonomy of RAG Privacy Risks and a RAG Privacy Process Diagram. Our work contributes to the study of privacy in RAG not only by conducting the first systematization of risks and mitigations, but also by uncovering important considerations when mitigating privacy risks in RAG systems and assessing the current maturity of proposed mitigations.

隐私安全RAGLLM风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。