利用RAG系统透明性漏洞,黑盒攻击可诱导可信伪造内容被召回。
The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems
- 基于RAG透明机制,通过观察来源和表述方式构造高召回伪造文档。
- 攻击后系统性能显著下降,且伪造内容自然可信,用户难辨真伪。
- 无需内部访问,适合研究系统安全与对抗攻击的从业者。
随着检索增强生成(RAG)系统广泛应用,已有多种攻击方法被提出以降低其性能。然而,多数现有方法依赖不切实际的假设,即外部攻击者可访问内部组件如检索器。为解决此问题,本文提出一种基于RAG悖论的真实黑盒攻击:该悖论源于系统为增强信任而向用户公开检索到的文档及其来源,这使得攻击者能观察哪些来源被使用及信息如何表述,进而构造更易被检索到的污染文档并上传至对应来源。此外,由于RAG系统直接将检索内容提供给用户,这些文档不仅需可被检索,还需在语义上自然可信,以维持用户对结果的信任。与以往仅关注可检索性的方法不同,本攻击同时考虑可检索性与用户信任度。离线与在线实验均表明,该方法在无内部访问条件下显著降低系统性能,同时生成自然、可信的污染文档。
原文摘要 · Abstract (English)
With the growing adoption of retrieval-augmented generation (RAG) systems, various attack methods have been proposed to degrade their performance. However, most existing approaches rely on unrealistic assumptions in which external attackers have access to internal components such as the retriever. To address this issue, we introduce a realistic black-box attack based on the RAG paradox, a structural vulnerability arising from the system's effort to enhance trust by revealing both the retrieved documents and their sources to users. This transparency enables attackers to observe which sources are used and how information is phrased, allowing them to craft poisoned documents that are more likely to be retrieved and upload them to the identified sources. Moreover, as RAG systems directly provide retrieved content to users, these documents must not only be retrievable but also appear natural and credible to maintain user confidence in the search results. Unlike prior work that focuses solely on improving document retrievability, our attack method explicitly considers both retrievability and user trust in the retrieved content. Both offline and online experiments demonstrate that our method significantly degrades system performance without internal access, while generating natural-looking poisoned documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。