提出新型污染攻击,让误导信息看似合理,骗过问答系统的纠错机制。
PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates

- 将污染信息伪装成合理更新,避免与系统已有知识直接冲突。
- 在45组实验中35次成功诱导目标答案,平均成功率提升9.7个百分点。
- 适合研究大模型安全、对抗攻击的学者或工程师参考。
在检索增强生成(RAG)中,后检索冲突解决机制用于处理噪声或矛盾的检索内容。然而,该机制对知识污染攻击的鲁棒性尚未充分研究。现有黑盒污染方法均以与解决机制认定结论直接矛盾的方式主张目标答案,而这种矛盾恰恰是攻击被检测的信号。本文提出PURPOSE,一种严格的黑盒污染攻击,将注入内容重构为最小化冲突的更新,而非对立声明。PURPOSE提取与查询相关的事实,近似解析解决机制可能参考的内容,再基于这些事实构建一个枢纽事件,使注入内容与解决机制可验证的逻辑保持一致,同时引导生成器输出目标答案。在三个QA基准、五种生成器和三种冲突解决方法上,PURPOSE在45组设置中取得35次最高攻击成功率(ASR),且平均比最强基线高出9.7个百分点。结果表明,该污染方法能有效绕过RAG中的冲突解决机制,并揭示非矛盾式注入是一种实用的攻击模式。
原文摘要 · Abstract (English)
In Retrieval-Augmented Generation (RAG), post-retrieval conflict resolution arbitrates among noisy or contradictory retrieved passages. However, the robustness of this safeguard against knowledge poisoning has not been adequately studied. Existing black-box poisoning methods all assert the target answer in frontal contradiction with what the resolver treats as settled, the very signal these methods are built to detect. We propose PURPOSE, a strict black-box poisoning attack that reframes the injection as an update that minimizes conflict, rather than as a counter-claim. PURPOSE extracts query-related facts approximating the resolver's possible reference, then grounds a pivot event in them to keep the injection consistent with what the resolver might verify while steering the generator toward the target answer. Across three QA benchmarks, five generators, and three conflict-resolution methods, PURPOSE attains the highest attack success rate (ASR) in 35 of 45 settings and exceeds the strongest prior attack with +9.7 mean ASR points. These results show that our poisoning method is effective against conflict resolution in RAG and identify non-contradicting injection as a practical mode to enhance poisoning attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。