arXiv:2502.18290cs.CV2025-02CVPR被引 19

攻击自监督视觉编码器,诱导大模型产生隐蔽幻觉。

Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models

  • 通过优化触发器与后门学习,实现对自监督编码器的隐蔽攻击。
  • 攻击成功率超99%,导致视觉理解错误率上升77.6%。
  • 适用于安全敏感的多模态大模型,检测难度高。

自监督学习(SSL)视觉编码器可生成高质量图像表征,已成为大型视觉语言模型(LVLMs)视觉模态的关键组件。由于训练成本高昂,预训练编码器被广泛共享并部署于多个LVLM中,这些模型具有重要的安全性和社会意义。在此背景下,我们揭示了一种新型后门威胁:仅需攻破视觉编码器,即可在这些LVLM中诱发显著的视觉幻觉。由于编码器的共享与复用,多个下游LVLM可能继承后门行为,导致后门大规模扩散。本文提出BadVision,首个针对此类编码器的后门攻击方法,结合新颖的触发器优化与后门学习技术。我们在两种类型SSL编码器及多个LVLM上,跨八个基准测试评估了该方法。结果表明,BadVision能以超过99%的成功率引导LVLM产生攻击者指定的幻觉,造成77.6%的相对视觉理解误差,同时保持高度隐蔽性。现有最先进后门检测方法无法有效识别此攻击。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) vision encoders learn high-quality image representations and thus have become a vital part of developing vision modality of large vision language models (LVLMs). Due to the high cost of training such encoders, pre-trained encoders are widely shared and deployed into many LVLMs, which are security-critical or bear societal significance. Under this practical scenario, we reveal a new backdoor threat that significant visual hallucinations can be induced into these LVLMs by merely compromising vision encoders. Because of the sharing and reuse of these encoders, many downstream LVLMs may inherit backdoor behaviors from encoders, leading to widespread backdoors. In this work, we propose BadVision, the first method to exploit this vulnerability in SSL vision encoders for LVLMs with novel trigger optimization and backdoor learning techniques. We evaluate BadVision on two types of SSL encoders and LVLMs across eight benchmarks. We show that BadVision effectively drives the LVLMs to attacker-chosen hallucination with over 99% attack success rate, causing a 77.6% relative visual understanding error while maintaining the stealthiness. SoTA backdoor detection methods cannot detect our attack effectively.

后门攻击视觉编码器大模型安全自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。