构建首个覆盖49类图像扰动的VLM鲁棒性评测基准
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
- 设计133种带严重度分级的图像畸变测试场景
- 发现空间畸变比视觉明显畸变更影响模型表现,降幅最高达34个百分点
- 揭示当前VLM对重采样与几何变换敏感,适合评估模型鲁棒性
视觉语言模型(VLMs)在高质量标准数据集上表现优异,但我们对其在真实世界图像畸变下的性能仍缺乏充分理解。本文提出VLM-RobustBench,涵盖噪声、模糊、天气、数字和几何扰动共49种增强类型,按低/中/高严重度及二值化变换组合形成133个受损测试场景。我们在两个互补基准MMBench(视觉对齐)和MMMU-Pro(推理导向)上评估了Qwen、InternVL、Molmo、Gemma四个系列模型。结果表明,视觉严重度并非难度的可靠预测指标:低严重度的空间扰动(如glass_blur)平均使MMBench准确率下降约8个百分点,而重采样与几何畸变(如upsample、elastic_transform)导致的最大降幅可达34个百分点。总体显示当前VLM语义能力强但空间结构脆弱,亟需建立强调重采样与几何不变性的新型鲁棒性评估与训练范式。
原文摘要 · Abstract (English)
Vision-language models (VLMs) achieve strong performance on standard, high-quality datasets, but we still do not fully understand how they perform under real-world image distortions. We present VLM-RobustBench, a benchmark spanning 49 augmentation types across noise, blur, weather, digital, and geometric perturbations, evaluated under graded severities (low/mid/high) and binary transforms, yielding 133 corrupted settings. We evaluate VLMs from four families (Qwen, InternVL, Molmo, Gemma) on two complementary benchmarks: MMBench (visually grounded) and MMMU-Pro (reasoning-oriented). Our results reveal that visual severity is a weak predictor of difficulty: low-severity spatial perturbations often degrade performance more than visually severe photometric corruptions. In particular, low-severity glass_blur reduces MMBench accuracy by about 8 pp on average across models, while the largest drops arise from resampling and geometric distortions (e.g., upsample, elastic_transform), reaching up to 34 pp. Overall, our findings suggest current VLMs are semantically strong but spatially fragile, motivating the definition of novel robustness evaluation protocols and training regimes that emphasize resampling and geometric invariances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。