频域扰动可让视觉语言模型误判图像真伪与内容,暴露其脆弱性。
On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- 在频域对图像进行微小结构扰动,改变VLM输出结果。
- 5种主流VLM在10个数据集上均出现判断失误,准确率显著下降。
- 适用于评估AI内容检测系统可靠性,尤其关注视觉语言模型部署场景。
视觉语言模型(VLMs)被广泛用于图像内容理解,如自动生成描述和深度伪造检测。本文揭示了当这些模型面对频域中的细微、有结构的扰动时,存在关键漏洞。我们设计了针对频域的图像变换方法,能系统性地干扰真实与合成图像下VLM的输出。实验表明,该扰动方法在五种前沿VLM(包括不同参数量的Qwen2/2.5和BLIP系列)上具有通用性;在十个真实与生成图像数据集上,模型判断高度依赖频率线索而非语义内容。关键发现是:肉眼不可见的空间频域变换即可暴露部署于自动图像描述与真实性检测任务中VLM的脆弱性。研究在黑盒条件下挑战了现有VLM的可靠性,强调亟需构建更鲁棒的多模态感知系统。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly used as perceptual modules for visual content reasoning, including through captioning and DeepFake detection. In this work, we expose a critical vulnerability of VLMs when exposed to subtle, structured perturbations in the frequency domain. Specifically, we highlight how these feature transformations undermine authenticity/DeepFake detection and automated image captioning tasks. We design targeted image transformations, operating in the frequency domain to systematically adjust VLM outputs when exposed to frequency-perturbed real and synthetic images. We demonstrate that the perturbation injection method generalizes across five state-of-the-art VLMs which includes different-parameter Qwen2/2.5 and BLIP models. Experimenting across ten real and generated image datasets reveals that VLM judgments are sensitive to frequency-based cues and may not wholly align with semantic content. Crucially, we show that visually-imperceptible spatial frequency transformations expose the fragility of VLMs deployed for automated image captioning and authenticity detection tasks. Our findings under realistic, black-box constraints challenge the reliability of VLMs, underscoring the need for robust multimodal perception systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。