发现视觉问答模型依赖文本捷径,提出新方法评估并过滤虚假视觉依赖。
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

- 引入盲测差距与视觉增益量化模型对视觉信息的依赖程度。
- 在MM-AU数据集上移除视频反而提升准确率,证明存在严重文本捷径。
- 提出实例级快捷方式分数,无需训练即可筛选易受干扰的问题。
高基准准确率未必代表模型真正使用了视觉证据。我们研究交通事故视频问答(VideoQA)中的这一问题,正确答案应依赖场景特定的视觉信息,但模型可能仅通过文本捷径推断。通过对四个公开基准的审计发现,多个近期开放权重视觉-语言模型(VLMs)在无视频输入时仍能保持竞争力,甚至表现更优。在MM-AU基准上,移除视频持续提升准确率,增加帧数反而导致性能下降。为此,我们提出两个数据集级诊断指标:盲测差距(Blind Gap),衡量文本单独的超随机表现;视觉增益(Visual Gain),衡量添加视频的边际收益。进一步提出实例级快捷方式分数,结合文本置信度与视觉必要性信号,实现无需训练的连续过滤,剔除易受捷径影响的问题。由此生成的子集降低了捷径偏差,提升了视觉对齐能力。研究揭示不同基准间视觉对齐质量差异显著,强调在安全关键型VideoQA中,仅追求高准确率不足,必须进行视觉对齐评估。
原文摘要 · Abstract (English)
High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers should depend on scene-specific visual evidence but may instead be inferred from textual shortcuts. Through an audit of four public benchmarks, we find that several recent open-weight Vision-Language Models (VLMs) perform competitively, and sometimes better, without video input. On the MM-AU benchmark, removing video consistently improves accuracy, and adding more frames further degrades performance. To quantify visual dependence, we introduce two dataset-level diagnostics: Blind Gap, measuring above-chance text-only performance, and Visual Gain, measuring the marginal benefit of adding video. We further propose an instance-level Shortcut Score that combines text-only confidence with visual necessity signals, enabling continuous, training-free filtering of shortcut-prone questions. The resulting subsets reduce shortcut bias and improve visual grounding. Our findings reveal large differences in grounding quality across benchmarks and show that visually grounded evaluation, not just high accuracy, is essential in safety-critical VideoQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。