用真实雨天特征攻击视觉语言模型,暴露其鲁棒性缺陷。
A Semantic Decoupling-Based Two-Stage Rainy-Day Attack for Revealing Weather Robustness Deficiencies in Vision-Language Models
- 分两阶段建模雨天影响:先全局调制语义空间,再精细模拟雨滴与光照变化。
- 即使约束严格,也能引发主流模型严重语义错配。
- 适合关注模型安全与真实场景鲁棒性的研究者。
视觉语言模型(VLMs)在标准视觉条件下训练,表现优异,但对真实天气条件的鲁棒性及跨模态语义对齐的稳定性仍研究不足。本文聚焦雨天场景,提出首个基于语义解耦的两阶段对抗框架,利用真实天气攻击VLMs。第一阶段通过低维全局调制改变嵌入空间,逐步弱化原始语义决策边界;第二阶段显式建模多尺度雨滴外观与降雨引起的光照变化,优化不可导的天气空间以诱导稳定语义偏移。该框架在非像素参数空间生成物理可信且可解释的扰动。多任务实验表明,即使高度约束的物理合理天气扰动,也能在主流VLMs中引发显著语义错配,威胁实际部署的安全性。消融实验证明光照建模和多尺度雨滴结构是关键驱动因素。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are trained on image-text pairs collected under canonical visual conditions and achieve strong performance on multimodal tasks. However, their robustness to real-world weather conditions, and the stability of cross-modal semantic alignment under such structured perturbations, remain insufficiently studied. In this paper, we focus on rainy scenarios and introduce the first adversarial framework that exploits realistic weather to attack VLMs, using a two-stage, parameterized perturbation model based on semantic decoupling to analyze rain-induced shifts in decision-making. In Stage 1, we model the global effects of rainfall by applying a low-dimensional global modulation to condition the embedding space and gradually weaken the original semantic decision boundaries. In Stage 2, we introduce structured rain variations by explicitly modeling multi-scale raindrop appearance and rainfall-induced illumination changes, and optimize the resulting non-differentiable weather space to induce stable semantic shifts. Operating in a non-pixel parameter space, our framework generates perturbations that are both physically grounded and interpretable. Experiments across multiple tasks show that even physically plausible, highly constrained weather perturbations can induce substantial semantic misalignment in mainstream VLMs, posing potential safety and reliability risks in real-world deployment. Ablations further confirm that illumination modeling and multi-scale raindrop structures are key drivers of these semantic shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。