提出曲边分形对抗贴纸,可骗过可见光-红外视觉语言模型。
Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

- 用贝塞尔曲线替代直边,生成多尺度自相似曲边分形贴纸。
- 在可见光与红外图像中注入螺旋纹理干扰,提升攻击效果。
- 生成的对抗样本能跨任务迁移,适用于分类、描述、问答等任务。
视觉语言模型(VLMs)在多种多模态任务中表现强劲,但在可见光-红外(VIS-IR)场景下的对抗鲁棒性尚未充分研究。这一空白尤为关键,因为VIS-IR感知广泛应用于复杂成像条件下的真实感知系统。为此,本文提出CFGPatch——一种基于三角分形几何的曲边分形对抗贴纸框架,用于攻击VIS-IR VLMs。CFGPatch以贝塞尔曲线取代传统刚性直边,保留多尺度分形自相似性的同时,实现更平滑的轮廓、更丰富的方向变化和更强的形状可变形性。同时,设计了模态特定的弗雷泽螺旋渲染机制,在可见光与红外图像中引入细粒度纹理畸变与误导性感知线索。通过全局曲边分形结构与局部螺旋式外观干扰的结合,有效破坏模型对形状与纹理的识别。进一步采用期望变换(EOT)增强对常见图像变换的鲁棒性。大量实验表明,CFGPatch显著提升了攻击效果与鲁棒性,优于标准贴纸基线。此外,针对零样本分类优化的对抗样本在图像描述与视觉问答任务中也表现出强跨任务迁移能力,验证了其广泛泛化性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, but their adversarial robustness in visible-infrared (VIS-IR) scenarios remains underexplored. This gap is critical because VIS-IR sensing is widely used in real-world perception systems to support reliable understanding under challenging imaging conditions. To address this cross-modal threat setting, we propose CFGPatch, a curved-edge fractal geometric adversarial patch framework for attacking VIS-IR VLMs. CFGPatch builds on triangular fractal geometry and replaces rigid straight-edged primitives with Bezier-curved elements, preserving multi-scale fractal self-similarity while introducing smoother contours, richer directional variation, and more flexible shape deformation. In addition, we design a modality-specific Fraser-spiral rendering mechanism to inject fine-grained texture distortions and misleading perceptual cues into visible and infrared images. By coupling global curved-fractal geometry with local spiral-based appearance interference, CFGPatch disrupts both shape perception and texture interpretation. We further adopt expectation over transformation (EOT) to improve robustness against common image-level transformations. Extensive experiments show that CFGPatch effectively fools VIS-IR VLMs and consistently outperforms standard patch baselines in attack effectiveness and robustness. Moreover, adversarial samples optimized for zero-shot classification transfer well to image captioning and visual question answering, demonstrating strong cross-task transferability and generalizability across downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。