arXiv:2412.03002cs.CV2024-12ICCV被引 5

提出首个生成对抗性3D变换样本的框架,评估视觉语言模型在真实3D变化下的鲁棒性。

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?

  • 基于单视角生成可复现的对抗性3D变换样本,利用生成式3D先验实现零样本姿态操控。
  • 设计自然度奖励模型,确保对抗样本视觉真实,避免幻觉或不自然内容。
  • 构建首个针对真实3D变化的VQA评测基准MM3DTBench,适用于多类VLM架构与任务。

视觉语言模型(VLMs)展现出卓越的泛化能力,但在动态真实场景中的鲁棒性仍鲜有研究。为系统评估其对真实世界3D变化的鲁棒性,我们提出AdvDreamer——首个能从单视图观测生成物理可复现对抗性3D变换(Adv-3DT)样本的框架。AdvDreamer包含三项核心创新:首先,为在缺乏先验知识的情况下精准刻画真实3D变化,设计基于生成式3D先验的零样本单目姿态操控流程;其次,为确保最恶劣情况下的对抗样本视觉质量,提出自然度奖励模型,在对抗优化中提供连续自然度正则化,有效防止收敛至幻觉或不自然元素;第三,为支持跨多种VLM架构与视觉-语言任务的系统评估,引入逆语义概率损失作为对抗优化目标,仅在基础的视觉-文本对齐空间中操作。基于生成的高攻击性与强迁移性的Adv-3DT样本,我们构建了首个面向挑战性3D变化的视觉问答基准数据集MM3DTBench。对代表性VLMs的广泛评估显示,真实世界3D变化会对各类任务中的模型性能造成严重威胁。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have exhibited remarkable generalization capabilities, yet their robustness in dynamic real-world scenarios remains largely unexplored. To systematically evaluate VLMs' robustness to real-world 3D variations, we propose AdvDreamer, the first framework capable of generating physically reproducible Adversarial 3D Transformation (Adv-3DT) samples from single-view observations. In AdvDreamer, we integrate three key innovations: Firstly, to characterize real-world 3D variations with limited prior knowledge precisely, we design a zero-shot Monocular Pose Manipulation pipeline built upon generative 3D priors. Secondly, to ensure the visual quality of worst-case Adv-3DT samples, we propose a Naturalness Reward Model that provides continuous naturalness regularization during adversarial optimization, effectively preventing convergence to hallucinated or unnatural elements. Thirdly, to enable systematic evaluation across diverse VLM architectures and visual-language tasks, we introduce the Inverse Semantic Probability loss as the adversarial optimization objective, which solely operates in the fundamental visual-textual alignment space. Based on the captured Adv-3DT samples with high aggressiveness and transferability, we establish MM3DTBench, the first VQA benchmark dataset tailored to evaluate VLM robustness under challenging 3D variations. Extensive evaluations of representative VLMs with varying architectures reveal that real-world 3D variations can pose severe threats to model performance across various tasks.

视觉语言模型3D变化对抗攻击鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。