研究多个对手同时攻击检索增强生成系统时的对抗策略与效果。
Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation
- 提出竞争性攻击新范式,模拟多方对手在相同查询上争夺控制权。
- 发现多数单敌攻击策略在多方竞争下失效,甚至反向效果。
- 构建标准评测框架PoisonArena,支持多敌手真实场景下的攻防测试。
检索增强生成(RAG)系统虽提升了大语言模型的事实准确性,但仍易受检索污染攻击,即攻击者在知识库中植入篡改内容。以往研究多基于单一攻击者假设,但实际中高价值或高曝光查询常吸引多个目标冲突的攻击者。本文基于真实案例,提出竞争性攻击设定:多个攻击者同时尝试将同一或相近查询引导至不同目标。我们形式化该威胁模型,并引入竞争有效性作为衡量攻击者在竞争中优势的指标。大量实验表明,许多在单攻击者环境下有效的策略在多敌手条件下显著退化,出现性能反转,暴露出传统指标(如成功率、F1)的局限性。此外,我们提出了PoisonArena——一个标准化的评估框架与基准,用于在多攻击者的真实场景下评测毒化攻击与防御方法。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems improve the factual grounding of large language models (LLMs) but remain vulnerable to retrieval poisoning, where adversaries seed the corpus with manipulated content. Prior work largely evaluates this threat under a simplified single-attacker assumption. In practice, however, high-value or high-visibility queries attract multiple adversaries with conflicting objectives. Motivated by real cases, we introduce the setting of competing attacks, in which multiple attackers simultaneously attempt to steer the same or closely related query toward different targets. We formalize this threat model and propose competitive effectiveness, a metric that quantifies an attacker's advantage under competition. Extensive experiments show that many strategies that succeed in the single-attacker regime degrade markedly under competition, revealing performance inversions and highlighting the limits of conventional metrics such as attack success rate and F1. Furthermore, we present PoisonArena, a standardized framework and benchmark for evaluating poisoning attacks and defenses under realistic, multi-adversary conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。