ALDEN通过主动学习与分布估计,高效提取RAG系统中的私有数据
ALDEN: Boosting Private Data Extraction from Retrieval-Augmented Generation Systems via Active Learning and Distribution Estimation

- 利用主动学习生成多样化恶意查询,提升攻击覆盖面
- 基于知识库分布设计动态算法,使查询更精准指向私有数据
- 在多个测试中显著超越现有攻击方法,适合安全研究者参考
检索增强生成(RAG)广泛用于通过外部知识检索提升大语言模型的可靠性与泛化能力。然而,近期研究发现RAG系统仍易受数据提取攻击:攻击者可通过在用户查询中嵌入恶意指令,窃取私有数据。尽管此类攻击可行,但现有方法通常提取率低、实用性有限。本文提出ALDEN,一种新型攻击方法,能高效、有效地从RAG系统中提取私有数据。首先,采用主动学习策略生成多样化的恶意查询,以提高数据提取率;其次,观察到底层知识库的数据分布可为查询生成提供关键指导,因此引入基于衰减的动态算法来估计主题分布。通过结合二者,ALDEN在全面评估中显著优于当前最先进的方法。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) is widely used to augment large language models with external knowledge retrieval to improve reliability and generalization. However, recent studies have shown that RAG systems remain vulnerable to data extraction attacks, where adversaries can extract private data by embedding malicious commands into user queries. Despite their feasibility, existing attacks typically suffer from low data extraction rates and limited practical effectiveness. Here, we propose ALDEN, a novel attack that effectively and efficiently extracts private data from RAGs. First, we employ active learning to diversify malicious queries and improve data extraction rates. Second, we observe that the data distribution of the underlying knowledge base provides valuable guidance for query generation and introduce a decay-based dynamic algorithm to estimate the corresponding topic distribution. By combining them together, we demonstrate that ALDEN substantially outperforms state-of-the-art methods through comprehensive evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。