arXiv:2503.06903cs.CV2025-03ICCV被引 12

提出光照攻击框架,暴露视觉语言模型在光照变化下的脆弱性

When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack

  • 将全局光照分解为可调控的点光源,实现精细光照控制
  • 使先进模型如LLaVA-1.6性能显著下降,攻击自然度高
  • 适合研究模型鲁棒性或对抗攻击的学者参考

视觉语言模型(VLM)在诸多任务中表现优异,但其对真实世界光照变化的鲁棒性仍缺乏系统研究。为此,我们提出首个系统评估VLM光照鲁棒性的框架——光照变换攻击(ITA)。针对两大挑战:(1)如何建模具有细粒度控制的全局光照以生成多样光照条件;(2)如何在保持自然性的同时确保攻击有效性。我们创新性地基于光照渲染方程将全局光照分解为多个参数化点光源,从而实现以往方法无法捕捉的多样化光照变化。结合物理光照重建技术,可精确还原原场景中的光照交互,实现细粒度控制。针对第二挑战,通过在光照重建模型的潜在空间中操控光照而非直接像素操作,天然保留物理光照先验;同时引入感知约束保证与原图视觉一致性,多样性约束防止光源聚集。大量实验表明,ITA能显著降低LLaVA-1.6等先进VLM性能,同时保持良好自然度,揭示了视觉语言模型在光照变化下的关键脆弱性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable success in various tasks, yet their robustness to real-world illumination variations remains largely unexplored. To bridge this gap, we propose \textbf{I}llumination \textbf{T}ransformation \textbf{A}ttack (\textbf{ITA}), the first framework to systematically assess VLMs' robustness against illumination changes. However, there still exist two key challenges: (1) how to model global illumination with fine-grained control to achieve diverse lighting conditions and (2) how to ensure adversarial effectiveness while maintaining naturalness. To address the first challenge, we innovatively decompose global illumination into multiple parameterized point light sources based on the illumination rendering equation. This design enables us to model more diverse lighting variations that previous methods could not capture. Then, by integrating these parameterized lighting variations with physics-based lighting reconstruction techniques, we could precisely render such light interactions in the original scenes, finally meeting the goal of fine-grained lighting control. For the second challenge, by controlling illumination through the lighting reconstrution model's latent space rather than direct pixel manipulation, we inherently preserve physical lighting priors. Furthermore, to prevent potential reconstruction artifacts, we design additional perceptual constraints for maintaining visual consistency with original images and diversity constraints for avoiding light source convergence. Extensive experiments demonstrate that our ITA could significantly reduce the performance of advanced VLMs, e.g., LLaVA-1.6, while possessing competitive naturalness, exposing VLMS' critical illuminiation vulnerabilities.

光照攻击视觉语言模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。