arXiv:2411.15673cs.CV2024-11CVPR被引 11

通过外部知识对齐,提升视觉语言模型抗后门与投毒攻击能力

Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment

  • 利用语言模型提取外部知识,约束视觉注意力与知识对齐程度
  • 在多个数据集和模型上防御成功率超90%,且不影响模型性能
  • 无需推理修改,适合实际部署的视觉语言系统

近年来,基于自监督目标训练的视觉语言模型受到广泛关注。然而,大规模网络爬取数据的使用使其面临后门攻击和投毒攻击等安全威胁。本文提出一种针对对比学习型视觉语言模型的防御方法,利用语言模型提取的外部知识,阻止模型学习那些与外部知识弱对齐的图像区域之间的关联。具体通过施加约束,使模型对视觉区域的关注度与其与外部知识的对齐程度成正比。我们在多个数据集和多种近期攻击方法下进行了大量实验,结果表明,该方法在不同设置下均能有效防御攻击,同时保持模型性能,且推理阶段无需任何修改。

原文摘要 · Abstract (English)

In recent years there has been enormous interest in vision-language models trained using self-supervised objectives. However, the use of large-scale datasets scraped from the web for training also makes these models vulnerable to potential security threats, such as backdooring and poisoning attacks. In this paper, we propose a method for mitigating such attacks on contrastively trained vision-language models. Our approach leverages external knowledge extracted from a language model to prevent models from learning correlations between image regions which lack strong alignment with external knowledge. We do this by imposing constraints to enforce that attention paid by the model to visual regions is proportional to the alignment of those regions with external knowledge. We conduct extensive experiments using a variety of recent backdooring and poisoning attacks on multiple datasets and architectures. Our results clearly demonstrate that our proposed approach is highly effective at defending against such attacks across multiple settings, while maintaining model utility and without requiring any changes at inference time

视觉语言模型安全防御知识对齐对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。