arXiv:2609.07048cs.CV2026-09

在视觉语言模型中植入频域隐蔽后门,攻击成功率超99%。

FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models

论文配图:FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models
图 1 · 摘自论文原文
  • 通过混合频谱幅度与原图相位生成不可见触发器。
  • 在Flickr8k上对三模型攻击成功率最高达99.8%。
  • 适合研究模型安全的人员关注,规避传统可见触发器。

视觉语言模型(VLMs)在开放式图文生成任务中表现优异,但其多模态特性使其易受后门攻击。现有后门触发器多为空间、文本或双模态形式,可能产生局部或可识别的模式。本文探索新攻击面,提出 extsc{FreqDoor},一种训练时的频域后门攻击方法。该方法选择性融合触发源图像的幅值谱成分,同时保留干净图像的相位,生成空间分布且视觉上不可察觉的触发器,不修改文本输入。我们在BLIP-2、InstructBLIP和LLaVA上评估该攻击,针对图像描述和视觉问答任务。在Flickr8k数据集上,三模型攻击成功率分别为99.6%、99.8%和98.4%,同时保持生成句义质量;在VQAv2上,对应成功率为99.6%、92.4%和79.6%。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. However, their multimodal nature makes them persistently vulnerable to backdoor attacks. Existing backdoor triggers for VLMs are either spatial, textual, or bimodal, which may yield localized or recognizable trigger patterns. In this work, we explore a different attack surface and propose \ textsc {FreqDoor}, a training-time backdoor attack that implants triggers in the frequency domain. \ textsc {FreqDoor} mixes amplitude-spectrum components from a trigger-source image selectively while preserving the phase of a clean image to generate a spatially distributed and visually imperceptible trigger without modifying the textual input. We evaluate the attack on BLIP-2, InstructBLIP, and LLaVA for image captioning and visual question answering. On Flickr8k, \ textsc {FreqDoor} achieves attack success rates of $99.6\%$, $99.8\%$, and $98.4\%$ on the three models, respectively, while preserving the semantic quality of the generated captions. On VQAv2, the corresponding attack success rates are $99.6\%$, $92.4\%$, and $79.6\%$.

后门攻击频域视觉语言模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。