首个全面评估RAG系统中毒攻击的基准框架,揭示现有防御失效风险。
Benchmarking Poisoning Attacks against Retrieval-Augmented Generation
- 构建涵盖13种攻击与7种防御的综合评测体系
- 扩展数据集上攻击成功率下降超50%,暴露防御短板
- 多类RAG架构均易受攻击,适合安全研究者参考
检索增强生成(RAG)通过引入外部知识有效缓解大模型幻觉问题,但其集成也带来新的安全威胁,尤其是中毒攻击。尽管已有研究探索多种攻击策略,但对其实用威胁的系统性评估仍属空白。为此,本文提出首个针对RAG中毒攻击的综合性基准框架,覆盖5个标准问答数据集及10个扩展变体,包含13种攻击方法与7种防御机制,涵盖当前主流技术。基于该框架,我们对所有攻击与防御在完整数据谱系上进行了全面评估。结果表明:虽然现有攻击在标准数据集上表现良好,但在扩展版本上效果显著下降;同时,多种先进RAG架构——包括序列、分支、条件、循环结构,多轮对话、多模态及基于LLM的RAG代理系统——均仍易受攻击。更关键的是,当前防御手段难以提供可靠保护,凸显亟需更具鲁棒性和普适性的防御方案。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) has proven effective in mitigating hallucinations in large language models by incorporating external knowledge during inference. However, this integration introduces new security vulnerabilities, particularly to poisoning attacks. Although prior work has explored various poisoning strategies, a thorough assessment of their practical threat to RAG systems remains missing. To address this gap, we propose the first comprehensive benchmark framework for evaluating poisoning attacks on RAG. Our benchmark covers 5 standard question answering (QA) datasets and 10 expanded variants, along with 13 poisoning attack methods and 7 defense mechanisms, representing a broad spectrum of existing techniques. Using this benchmark, we conduct a comprehensive evaluation of all included attacks and defenses across the full dataset spectrum. Our findings show that while existing attacks perform well on standard QA datasets, their effectiveness drops significantly on the expanded versions. Moreover, our results demonstrate that various advanced RAG architectures, such as sequential, branching, conditional, and loop RAG, as well as multi-turn conversational RAG, multimodal RAG systems, and RAG-based LLM agent systems, remain susceptible to poisoning attacks. Notably, current defense techniques fail to provide robust protection, underscoring the pressing need for more resilient and generalizable defense strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。