用布料褶皱机制生成逼真攻击,让视觉语言模型看错图。
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models
- 基于三维布料褶皱机理,生成多尺度逼真非刚性扰动。
- 在零样本分类任务上优化后,攻击效果在图文生成任务中仍有效。
- 对多种先进视觉语言模型均有效,适合研究模型鲁棒性者参考。
视觉语言模型(VLMs)在零样本分类、图像描述和视觉问答等任务中表现出卓越的跨模态理解能力,但其对物理上合理的非刚性形变(如柔性表面褶皱)的鲁棒性仍不明确。本文提出一种受三维织物褶皱力学启发的参数化结构扰动方法,通过构建多尺度褶皱场,并结合位移场扭曲与表面一致的外观变化,生成逼真非刚性扰动。为平衡视觉自然度与攻击有效性,在低维参数空间设计分层适应度函数,并采用基于优化的搜索策略。评估采用两阶段框架:先在零样本分类代理任务上优化扰动,再测试其在生成任务上的可迁移性。实验表明,该方法显著降低多种先进VLM性能,且在图像描述和视觉问答任务中持续优于基线。
原文摘要 · Abstract (English)
Visual-Language Models (VLMs) have demonstrated exceptional cross-modal understanding across various tasks, including zero-shot classification, image captioning, and visual question answering. However, their robustness to physically plausible non-rigid deformations-such as wrinkles on flexible surfaces-remains poorly understood. In this work, we propose a parametric structural perturbation method inspired by the mechanics of three-dimensional fabric wrinkles. Specifically, our method generates photorealistic non-rigid perturbations by constructing multi-scale wrinkle fields and integrating displacement field distortion with surface-consistent appearance variations. To achieve an optimal balance between visual naturalness and adversarial effectiveness, we design a hierarchical fitness function in a low-dimensional parameter space and employ an optimization-based search strategy. We evaluate our approach using a two-stage framework: perturbations are first optimized on a zero-shot classification proxy task and subsequently assessed for transferability on generative tasks. Experimental results demonstrate that our method significantly degrades the performance of various state-of-the-art VLMs, consistently outperforming baselines in both image captioning and visual question-answering tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。