提出新基准V-FAT,检测多模态模型是否真看图而非猜文本。
V-FAT: Benchmarking Visual Fidelity Against Text-bias
- 构建三层次冲突测试框架,分离内部与外部文本偏见。
- 12个前沿模型在高语言干扰下视觉表现显著下降。
- 引入视觉鲁棒性评分,避免靠猜测得分的假象。
近期多模态大模型在标准视觉推理任务中表现优异,但存在过度依赖语言捷径而缺乏真实视觉理解的问题,我们称之为文本偏见。本文研究视觉感知与语言先验之间的根本矛盾,将偏见来源分解为两类:源自预训练语料统计相关性的内部语料偏见,以及对齐过程中产生的外部指令偏见。为此,我们提出V-FAT(Visual Fidelity Against Text-bias)诊断基准,包含4026个跨六个语义领域的VQA实例。该基准采用三层评估框架,逐步增加视觉证据与文本信息间的冲突:(L1)异常图像引发的内部偏见,(L2)误导性指令引发的外部偏见,(L3)两者协同作用的综合偏见。我们引入视觉鲁棒性分数(VRS),用于惩罚“侥幸”的语言猜测,奖励真正的视觉理解能力。对12个前沿多模态大模型的评估显示,尽管这些模型在现有基准上表现良好,但在高语言主导情境下出现显著的视觉退化现象。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on standard visual reasoning benchmarks. However, there is growing concern that these models rely excessively on linguistic shortcuts rather than genuine visual grounding, a phenomenon we term Text Bias. In this paper, we investigate the fundamental tension between visual perception and linguistic priors. We decouple the sources of this bias into two dimensions: Internal Corpus Bias, stemming from statistical correlations in pretraining, and External Instruction Bias, arising from the alignment-induced tendency toward sycophancy. To quantify this effect, we introduce V-FAT (Visual Fidelity Against Text-bias), a diagnostic benchmark comprising 4,026 VQA instances across six semantic domains. V-FAT employs a Three-Level Evaluation Framework that systematically increases the conflict between visual evidence and textual information: (L1) internal bias from atypical images, (L2) external bias from misleading instructions, and (L3) synergistic bias where both coincide. We introduce the Visual Robustness Score (VRS), a metric designed to penalize "lucky" linguistic guesses and reward true visual fidelity. Our evaluation of 12 frontier MLLMs reveals that while models excel in existing benchmarks, they experience significant visual collapse under high linguistic dominance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。