arXiv:2602.17168cs.CV2026-02被引 1

提出新型多模态后门攻击方法,隐蔽性强且抗检测能力强。

BadCLIP++: Stealthy and Persistent Backdoors in Multimodal Contrastive Learning

  • 设计语义融合微触发器,隐藏在任务相关区域,保持数据统计特性。
  • 仅用0.3%污染数据即达99.99%攻击成功率,物理攻击仍达65.03%。
  • 适合研究后门防御机制的学者,尤其关注隐蔽性与持久性的场景。

针对多模态对比学习模型的后门攻击面临隐蔽性与持久性双重挑战。现有方法在强检测或持续微调下易失效,主要因(1)跨模态不一致暴露触发模式,及(2)低污染率下梯度稀释加速后门遗忘。本文提出统一框架BadCLIP++,解决上述问题。为提升隐蔽性,引入语义融合QR微触发器,嵌入任务相关区域附近,保留干净数据统计特性并生成紧凑触发分布;结合目标对齐子集选择,在低注入率下强化信号。为增强持久性,通过半径收缩与中心对齐稳定触发嵌入,利用曲率控制与弹性权重巩固稳定模型参数,将解维持在低曲率宽谷中以抵抗微调。首次提供理论分析表明,在信任区域内,干净微调与后门目标的梯度共向,攻击成功率退化有上界。实验显示,仅0.3%污染率下,数字攻击成功率高达99.99%,超越基线11.4个百分点;在19种防御下,攻击成功率仍超99.90%,干净准确率下降不足0.8%。物理攻击成功率达65.03%,且对水印移除防御具鲁棒性。

原文摘要 · Abstract (English)

Research on backdoor attacks against multimodal contrastive learning models faces two key challenges: stealthiness and persistence. Existing methods often fail under strong detection or continuous fine-tuning, largely due to (1) cross-modal inconsistency that exposes trigger patterns and (2) gradient dilution at low poisoning rates that accelerates backdoor forgetting. These coupled causes remain insufficiently modeled and addressed. We propose BadCLIP++, a unified framework that tackles both challenges. For stealthiness, we introduce a semantic-fusion QR micro-trigger that embeds imperceptible patterns near task-relevant regions, preserving clean-data statistics while producing compact trigger distributions. We further apply target-aligned subset selection to strengthen signals at low injection rates. For persistence, we stabilize trigger embeddings via radius shrinkage and centroid alignment, and stabilize model parameters through curvature control and elastic weight consolidation, maintaining solutions within a low-curvature wide basin resistant to fine-tuning. We also provide the first theoretical analysis showing that, within a trust region, gradients from clean fine-tuning and backdoor objectives are co-directional, yielding a non-increasing upper bound on attack success degradation. Experiments demonstrate that with only 0.3% poisoning, BadCLIP++ achieves 99.99% attack success rate (ASR) in digital settings, surpassing baselines by 11.4 points. Across nineteen defenses, ASR remains above 99.90% with less than 0.8% drop in clean accuracy. The method further attains 65.03% success in physical attacks and shows robustness against watermark removal defenses.

后门攻击多模态隐蔽性对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。