arXiv:2505.18583cs.IR2025-05ACL被引 11

用微小扰动让黑盒RAG系统误检文档,误导生成答案。

The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems

  • 基于强化学习动态优化攻击策略,兼顾相关性、生成效果与自然度。
  • 仅用极小文本扰动即能成功诱使RAG系统检索目标文档,误导回答。
  • 适用于测试RAG系统安全性,尤其关注隐蔽对抗攻击的防御研究者。

我们研究了针对检索增强生成(RAG)系统的对抗攻击,以识别其脆弱性。重点在于生成人类难以察觉的对抗样本,提出一种新型的不可察觉检索-生成攻击方法。该任务旨在通过微小扰动,使原本不在初始前k个候选文档中的目标文档被检索到,从而影响最终生成的答案。为此,我们提出了ReGENT框架,基于强化学习追踪攻击者与目标RAG之间的交互,并根据相关性-生成-自然度奖励持续优化攻击策略。在新构建的事实性和非事实性问答基准上的实验表明,ReGENT在使用极小的不可察觉文本扰动时,显著优于现有攻击方法,能有效误导RAG系统。

原文摘要 · Abstract (English)

We explore adversarial attacks against retrieval-augmented generation (RAG) systems to identify their vulnerabilities. We focus on generating human-imperceptible adversarial examples and introduce a novel imperceptible retrieve-to-generate attack against RAG. This task aims to find imperceptible perturbations that retrieve a target document, originally excluded from the initial top-$k$ candidate set, in order to influence the final answer generation. To address this task, we propose ReGENT, a reinforcement learning-based framework that tracks interactions between the attacker and the target RAG and continuously refines attack strategies based on relevance-generation-naturalness rewards. Experiments on newly constructed factual and non-factual question-answering benchmarks demonstrate that ReGENT significantly outperforms existing attack methods in misleading RAG systems with small imperceptible text perturbations.

对抗攻击RAG安全生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。