arXiv:2508.17222cs.CRcs.AI2025-08ACL被引 6

Graph RAG虽提升回答质量,却更易泄露实体关系数据。

Exposing Privacy Risks in Graph Retrieval-Augmented Generation

  • 设计针对性攻击,探测图结构RAG的隐私漏洞。
  • 相比原始文本泄露减少,实体与关系信息泄露显著增加。
  • 适合关注大模型隐私安全的研究者和开发者。

检索增强生成(RAG)通过引入外部知识提升大语言模型性能。图结构RAG利用图谱知识实现更连贯、上下文丰富的回答,但其结构化检索机制引入了新的隐私风险。本文系统研究图RAG的数据提取漏洞,设计并实施针对性攻击,验证其对原始文本及实体关系等结构化数据的泄露风险。结果表明:尽管图RAG减少了原始文本泄露,但对实体及其关系信息的泄露风险显著上升。同时探讨潜在防御策略。本工作揭示了图RAG特有的隐私挑战,为构建更安全系统提供依据。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) is a powerful technique for enhancing Large Language Models (LLMs) with external, up-to-date knowledge. Graph RAG has emerged as an advanced paradigm that leverages graph-based knowledge structures to provide more coherent and contextually rich answers. However, the move from plain document retrieval to structured graph traversal introduces new, under-explored privacy risks. This paper investigates the data extraction vulnerabilities of the Graph RAG systems. We design and execute tailored data extraction attacks to probe their susceptibility to leaking both raw text and structured data, such as entities and their relationships. Our findings reveal a critical trade-off: while Graph RAG systems may reduce raw text leakage, they are significantly more vulnerable to the extraction of structured entity and relationship information. We also explore potential defense mechanisms to mitigate these novel attack surfaces. This work provides a foundational analysis of the unique privacy challenges in Graph RAG and offers insights for building more secure systems.

图RAG隐私安全大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。