arXiv:2504.16727cs.CVcs.AI2025-04ACL被引 10

发现大模型对物体位置尺度变化极度脆弱,揭示其架构缺陷。

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward

  • 构建新基准V²R-Bench,自动生成视觉变化数据集并量化鲁棒性
  • 21个模型均在简单识别任务上表现远低于预期,最差下降70%
  • 暴露模型位置偏差和类人视觉阈值,适合关注模型可靠性研究者

大型视觉语言模型(LVLMs)在多类视觉语言任务中表现优异,但其对自然场景中因视角与环境变化导致的位置、尺度、方向及上下文等基本视觉变化的鲁棒性仍缺乏系统研究。为此,本文提出V²R-Bench,一个涵盖自动化数据生成与严谨评估指标的综合性基准框架。通过对21个主流LVLMs的广泛评测,发现即使在复杂任务中表现优秀的模型,在基础物体识别任务上也显著退化,部分任务性能下降达70%。有趣的是,这些模型表现出与有效感受野理论相悖的位置偏差,并呈现类人视觉敏锐度阈值。通过组件级分析框架与新型对齐特征可视化方法,揭示其根源在于流水线中的误差累积与多模态对齐不足。合成数据实验进一步表明,此类缺陷源于根本性架构局限,亟需未来模型设计进行结构性创新。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) excel in various vision-language tasks. Yet, their robustness to visual variations in position, scale, orientation, and context that objects in natural scenes inevitably exhibit due to changes in viewpoint and environment remains largely underexplored. To bridge this gap, we introduce V$^2$R-Bench, a comprehensive benchmark framework for evaluating Visual Variation Robustness of LVLMs, which encompasses automated evaluation dataset generation and principled metrics for thorough robustness assessment. Through extensive evaluation on 21 LVLMs, we reveal a surprising vulnerability to visual variations, in which even advanced models that excel at complex vision-language tasks significantly underperform on simple tasks such as object recognition. Interestingly, these models exhibit a distinct visual position bias that contradicts theories of effective receptive fields, and demonstrate a human-like visual acuity threshold. To identify the source of these vulnerabilities, we present a systematic framework for component-level analysis, featuring a novel visualization approach for aligned visual features. Results show that these vulnerabilities stem from error accumulation in the pipeline architecture and inadequate multimodal alignment. Complementary experiments with synthetic data further demonstrate that these limitations are fundamentally architectural deficiencies, scoring the need for architectural innovations in future LVLM designs.

视觉语言模型鲁棒性评测架构缺陷多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。