arXiv:2410.01264cs.CV2024-10ICLR被引 32

攻击者仅用异常数据就能劫持视觉语言模型,且不破坏原功能。

Backdooring Vision-Language Models with Out-Of-Distribution Data

  • 利用分布外数据实现无原始训练数据的后门注入。
  • 在图像描述与视觉问答任务中成功触发后门,原任务性能几乎不变。
  • 揭示多模态模型严重安全隐患,适合安全研究者关注。

视觉语言模型(VLM)将计算机视觉与大语言模型结合,可基于图像生成详细文本描述,其重要性日益凸显。然而,针对VLM的后门攻击安全性研究仍不充分。以往工作常假设攻击者可访问原始训练数据,这在现实中并不现实。本文提出更贴近实际的攻击场景:攻击者仅能使用分布外(OOD)数据。我们提出VLOOD方法,实现两个关键贡献:(1) 在复杂图文任务中成功实施后门攻击,同时最大限度减少中毒输入对原语义的干扰;(2) 提出无需原始训练数据即可注入后门的新技术。在图像描述和视觉问答任务上的评估验证了VLOOD的有效性,揭示了VLM存在重大安全漏洞,为未来多模态模型对抗复杂威胁的研究奠定基础。

原文摘要 · Abstract (English)

The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to generate detailed text descriptions from visual inputs. Despite their growing importance, the security of VLMs, particularly against backdoor attacks, is under explored. Moreover, prior works often assume attackers have access to the original training data, which is often unrealistic. In this paper, we address a more practical and challenging scenario where attackers must rely solely on Out-Of-Distribution (OOD) data. We introduce VLOOD (Backdooring Vision-Language Models with Out-of-Distribution Data), a novel approach with two key contributions: (1) demonstrating backdoor attacks on VLMs in complex image-to-text tasks while minimizing degradation of the original semantics under poisoned inputs, and (2) proposing innovative techniques for backdoor injection without requiring any access to the original training data. Our evaluation on image captioning and visual question answering (VQA) tasks confirms the effectiveness of VLOOD, revealing a critical security vulnerability in VLMs and laying the foundation for future research on securing multimodal models against sophisticated threats.

视觉语言模型后门攻击安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。