通过反馈引导跨模态对抗搜索,提升对视觉语言模型的攻击效果
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
- 设计模态互斥损失,让正确图文对远离,错误配对靠近
- 利用目标模型反馈迭代优化对抗样本,突破现有攻击局限
- 首次在跨模态场景中探索目标模型反馈的对抗边界,适合安全研究者
尽管视觉-语言预训练(VLP)模型在跨模态任务上取得了显著进展,但仍易受对抗攻击。基于数据增强和跨模态交互生成可迁移对抗样本的迁移式黑盒攻击已成为主流方法,因其更贴近真实场景。然而,由于不同模型间特征表示差异,其迁移性可能受限。为此,本文提出一种新攻击范式——反馈式模态互斥搜索(FMMS)。FMMS引入新型模态互斥损失(MML),旨在使匹配的图像-文本对在特征空间中分离,同时随机拉近不匹配对的距离,从而指导对抗样本的更新方向。此外,FMMS利用目标模型反馈迭代精炼对抗样本,推动其进入对抗区域。据我们所知,这是首个利用目标模型反馈探索多模态对抗边界的尝试。在Flickr30K和MSCOCO数据集上的图像-文本匹配任务上,大量实验证明FMMS显著优于当前最佳基线。
原文摘要 · Abstract (English)
Although vision-language pre-training (VLP) models have achieved remarkable progress on cross-modal tasks, they remain vulnerable to adversarial attacks. Using data augmentation and cross-modal interactions to generate transferable adversarial examples on surrogate models, transfer-based black-box attacks have become the mainstream methods in attacking VLP models, as they are more practical in real-world scenarios. However, their transferability may be limited due to the differences on feature representation across different models. To this end, we propose a new attack paradigm called Feedback-based Modal Mutual Search (FMMS). FMMS introduces a novel modal mutual loss (MML), aiming to push away the matched image-text pairs while randomly drawing mismatched pairs closer in feature space, guiding the update directions of the adversarial examples. Additionally, FMMS leverages the target model feedback to iteratively refine adversarial examples, driving them into the adversarial region. To our knowledge, this is the first work to exploit target model feedback to explore multi-modality adversarial boundaries. Extensive empirical evaluations on Flickr30K and MSCOCO datasets for image-text matching tasks show that FMMS significantly outperforms the state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。