让恶意内容在真实RAG系统中更难被发现且效果更强
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
- 用大模型自动生成能抗处理和查询变化的毒化内容
- 在多种数据集和配置下显著优于现有攻击方法
- 既可用于攻击测试,也能用于保护网页不被非法抓取
检索增强生成(RAG)在实际应用中越来越普遍,其基于参考内容的设计使输出显得可信。这种可信性促使研究者开发投毒攻击,通过将恶意内容注入知识源来操控RAG响应。然而,在真实RAG系统中,现有攻击效果严重下降。这源于两个被忽视的事实:(i) 内容在使用前常被处理,可能打碎毒化信息削弱效果;(ii) 用户查询往往与攻击设计时预设的不一致。这些因素导致从业者低估风险,产生虚假安全感。为此,我们提出Confundo,一种学习型投毒框架,通过微调大语言模型作为毒化生成器,实现高有效性、鲁棒性和隐蔽性。Confundo支持多种攻击目标,包括篡改事实正确性、诱导偏见观点和触发幻觉。通过解决上述挑战,Confundo在多个数据集和RAG配置下显著优于各类专用攻击,即使面对防御机制也表现优异。此外,我们还展示了其防御用途:可有效防止网页内容被未经授权地抓取并纳入RAG系统,且不影响用户体验。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) is increasingly deployed in real-world applications, where its reference-grounded design makes outputs appear trustworthy. This trust has spurred research on poisoning attacks that craft malicious content, inject it into knowledge sources, and manipulate RAG responses. However, when evaluated in practical RAG systems, existing attacks suffer from severely degraded effectiveness. This gap stems from two overlooked realities: (i) content is often processed before use, which can fragment the poison and weaken its effect, and (ii) users often do not issue the exact queries anticipated during attack design. These factors can lead practitioners to underestimate risks and develop a false sense of security. To better characterize the threat to practical systems, we present Confundo, a learning-to-poison framework that fine-tunes a large language model as a poison generator to achieve high effectiveness, robustness, and stealthiness. Confundo provides a unified framework supporting multiple attack objectives, demonstrated by manipulating factual correctness, inducing biased opinions, and triggering hallucinations. By addressing these overlooked challenges, Confundo consistently outperforms a wide range of purpose-built attacks across datasets and RAG configurations by large margins, even in the presence of defenses. Beyond exposing vulnerabilities, we also present a defensive use case that protects web content from unauthorized incorporation into RAG systems via scraping, with no impact on user experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。