首次对视觉语言模型发动后门攻击,触发特定文本输出。
TrojVLM: Backdoor Attack Against Vision Language Models
- 在图像输入中植入后门,使模型生成预设目标文本。
- 提出语义保持损失,确保原图内容不被破坏。
- 揭示多模态模型安全风险,适合安全研究者参考。
视觉语言模型(VLMs)将计算机视觉与大语言模型结合,可根据视觉输入生成详细文本描述,但随之带来新的安全漏洞。与以往针对单一模态或分类任务的研究不同,本文首次探索针对复杂图像到文本生成任务的后门攻击,提出TrojVLM。该方法在中毒图像触发时,可使模型输出预定的目标文本。同时,设计了一种新型语义保持损失,确保原始图像内容的语义完整性。在图像描述和视觉问答(VQA)任务上的评估验证了该方法的有效性:在保持原始语义不变的前提下,成功触发目标文本输出。本研究不仅揭示了VLMs在图像到文本生成中的关键安全风险,也为未来防御此类高级威胁提供了基础。
原文摘要 · Abstract (English)
The emergence of Vision Language Models (VLMs) is a significant advancement in integrating computer vision with Large Language Models (LLMs) to produce detailed text descriptions based on visual inputs, yet it introduces new security vulnerabilities. Unlike prior work that centered on single modalities or classification tasks, this study introduces TrojVLM, the first exploration of backdoor attacks aimed at VLMs engaged in complex image-to-text generation. Specifically, TrojVLM inserts predetermined target text into output text when encountering poisoned images. Moreover, a novel semantic preserving loss is proposed to ensure the semantic integrity of the original image content. Our evaluation on image captioning and visual question answering (VQA) tasks confirms the effectiveness of TrojVLM in maintaining original semantic content while triggering specific target text outputs. This study not only uncovers a critical security risk in VLMs and image-to-text generation but also sets a foundation for future research on securing multimodal models against such sophisticated threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。