arXiv:2410.02182cs.CVcs.CR2024-10被引 31

提出首个统一框架的隐形跨模态后门攻击,可隐蔽植入触发器并绕过防御。

BadCM: Invisible Backdoor Attack Against Cross-Modal Learning

  • 通过挖掘模态不变特征区域,定位隐蔽触发点
  • 在图文检索与VQA任务中攻击成功率超90%,且隐匿性强
  • 适用于多种模型,对现有防御手段有较强鲁棒性

尽管单模态学习取得显著进展,但跨模态学习中的后门攻击仍研究不足,主要因通用性差和隐蔽性弱。现有方法多沿用视觉单模态思路,难以应对多样化的跨模态场景,且难以生成不可感知的污染样本。本文提出一种新型双边后门攻击框架BadCM,通过跨模态挖掘策略识别模态不变成分作为污染目标区域,将精心设计的触发模式注入其中,使受害模型高效识别。该策略适配多种图像-文本跨模态模型,支持多类攻击场景。为提升隐蔽性,设计针对视觉与语言模态的专用生成器,将触发模式隐藏于模态不变区域。实验在跨模态检索与视觉问答任务上验证了方法的有效性与泛化能力,攻击成功率超过90%。同时,坏码能有效规避现有后门防御机制。代码已开源。

原文摘要 · Abstract (English)

Despite remarkable successes in unimodal learning tasks, backdoor attacks against cross-modal learning are still underexplored due to the limited generalization and inferior stealthiness when involving multiple modalities. Notably, since works in this area mainly inherit ideas from unimodal visual attacks, they struggle with dealing with diverse cross-modal attack circumstances and manipulating imperceptible trigger samples, which hinders their practicability in real-world applications. In this paper, we introduce a novel bilateral backdoor to fill in the missing pieces of the puzzle in the cross-modal backdoor and propose a generalized invisible backdoor framework against cross-modal learning (BadCM). Specifically, a cross-modal mining scheme is developed to capture the modality-invariant components as target poisoning areas, where well-designed trigger patterns injected into these regions can be efficiently recognized by the victim models. This strategy is adapted to different image-text cross-modal models, making our framework available to various attack scenarios. Furthermore, for generating poisoned samples of high stealthiness, we conceive modality-specific generators for visual and linguistic modalities that facilitate hiding explicit trigger patterns in modality-invariant regions. To the best of our knowledge, BadCM is the first invisible backdoor method deliberately designed for diverse cross-modal attacks within one unified framework. Comprehensive experimental evaluations on two typical applications, i.e., cross-modal retrieval and VQA, demonstrate the effectiveness and generalization of our method under multiple kinds of attack scenarios. Moreover, we show that BadCM can robustly evade existing backdoor defenses. Our code is available at https://github.com/xandery-geek/BadCM.

后门攻击跨模态隐形攻击安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。