攻击图像文本元数据可诱导多模态生成错误,91%成功率
Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation
- 仅篡改图片文字条目的元数据,不改动图像内容
- 在4个检索器和2个生成器上实现最高91%攻击成功率
- 现有防御策略对此类隐蔽攻击基本无效,适合安全研究者关注
检索增强生成(RAG)通过引入外部知识库提升多模态大模型的准确性并减少幻觉。然而,外部知识源也带来了新的攻击面。攻击者可注入恶意多模态内容,影响检索与下游生成。本文提出MM-MEPA,一种针对图像-文本条目元数据的隐蔽中毒攻击,仅修改元数据而保持视觉内容不变。该方法仍能引导多模态检索并诱发模型产生攻击者期望的响应。在多个基准设置下评估显示,MM-MEPA在4个检索器和2个多模态生成器上均达到高达91%的攻击成功率,显著破坏系统行为。我们还测试了代表性防御策略,发现其对这种仅操纵元数据的攻击几乎无效。研究揭示了多模态RAG中的关键漏洞,强调需发展更鲁棒、具备防御意识的检索与知识集成方法。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) has emerged as a powerful paradigm for enhancing multimodal large language models by grounding their responses in external, factual knowledge and thus mitigating hallucinations. However, the integration of externally sourced knowledge bases introduces a critical attack surface. Adversaries can inject malicious multimodal content capable of influencing both retrieval and downstream generation. In this work, we present MM-MEPA, a multimodal poisoning attack that targets the metadata components of image-text entries while leaving the associated visual content unaltered. By only manipulating the metadata, MM-MEPA can still steer multimodal retrieval and induce attacker-desired model responses. We evaluate the attack across multiple benchmark settings and demonstrate its severity. MM-MEPA achieves an attack success rate of up to 91\% consistently disrupting system behaviors across four retrievers and two multimodal generators. Additionally, we assess representative defense strategies and find them largely ineffective against this form of metadata-only poisoning. Our findings expose a critical vulnerability in multimodal RAG and underscore the urgent need for more robust, defense-aware retrieval and knowledge integration methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。