仅用极少量数据就能让事实核查系统被误导,且攻击效果极强。
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- 通过语义对齐的多注入技术,在不接触模型的情况下实现攻击。
- 在4个检索器、11个大模型上平均成功率86%,毒化率低至0.93×10⁻⁶。
- 适用于真实场景,即使有反证据仍能成功,适合安全与评测研究者。
知识投毒对检索增强生成(RAG)系统构成严重威胁,通过向知识库注入对抗性内容,诱使大语言模型(LLMs)生成受攻击者控制的、基于篡改上下文的输出。现有研究已揭示LLM对误导性或恶意检索内容的脆弱性。然而,在真实事实核查场景中,可信证据通常主导检索结果,更具挑战性。为此,本文将知识投毒扩展至事实核查场景,其中检索上下文包含真实支持或反驳证据。我们提出 extbf{ADMIT}( extbf{AD}versarial extbf{M}ulti- extbf{I}njection extbf{T}echnique),一种无需访问目标LLM、检索器或词级控制的少样本、语义对齐投毒攻击方法,可翻转事实判断并诱导欺骗性解释。大量实验表明,ADMIT在4个检索器、11个LLM和4个跨领域基准上均有效迁移,平均攻击成功率(ASR)达86%,毒化率仅为 $0.93 \times 10^{-6}$,即便存在强反证据仍具鲁棒性。相比现有最优攻击,其在所有设置下提升ASR达11.2%,暴露了真实RAG式事实核查系统的显著漏洞。
原文摘要 · Abstract (English)
Knowledge poisoning poses a critical threat to Retrieval-Augmented Generation (RAG) systems by injecting adversarial content into knowledge bases, tricking Large Language Models (LLMs) into producing attacker-controlled outputs grounded in manipulated context. Prior work highlights LLMs' susceptibility to misleading or malicious retrieved content. However, real-world fact-checking scenarios are more challenging, as credible evidence typically dominates the retrieval pool. To investigate this problem, we extend knowledge poisoning to the fact-checking setting, where retrieved context includes authentic supporting or refuting evidence. We propose \textbf{ADMIT} (\textbf{AD}versarial \textbf{M}ulti-\textbf{I}njection \textbf{T}echnique), a few-shot, semantically aligned poisoning attack that flips fact-checking decisions and induces deceptive justifications, all without access to the target LLMs, retrievers, or token-level control. Extensive experiments show that ADMIT transfers effectively across 4 retrievers, 11 LLMs, and 4 cross-domain benchmarks, achieving an average attack success rate (ASR) of 86\% at an extremely low poisoning rate of $0.93 \times 10^{-6}$, and remaining robust even in the presence of strong counter-evidence. Compared with prior state-of-the-art attacks, ADMIT improves ASR by 11.2\% across all settings, exposing significant vulnerabilities in real-world RAG-based fact-checking systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。