arXiv:2602.11244cs.CV2026-02被引 1

测试发现视频语言模型在时间与视觉理解上很脆弱,常出错却自信。

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

  • 设计五种受控压力测试,诊断模型弱点
  • 模型对倒放视频仍自信描述,忽视真实内容
  • 适合研究模型可靠性或评估视频理解能力者

本研究探讨视频语言模型(VidLMs)是否能稳健理解视频内容、时间顺序与运动。结果出人意料:多数模型表现不佳。我们提出REVEAL{}基准,通过五项受控压力测试评估模型在时间预期偏差、仅依赖语言捷径、迎合视频表象、相机运动敏感性以及时空遮挡鲁棒性方面的缺陷。测试涵盖主流开源与闭源模型,发现它们会自信描述倒放场景为正向,忽略视频内容回答问题,认同错误陈述,难以应对基础相机运动,且在简单时空遮挡下无法有效聚合信息。人类则轻松完成这些任务。同时,我们提供自动化数据生成流程,支持更广泛、可扩展的评估。我们将公开基准与代码,推动后续研究。

原文摘要 · Abstract (English)

This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, surprisingly, they often do not. We introduce REVEAL{}, a diagnostic benchmark that probes fundamental weaknesses of contemporary VidLMs through five controlled stress tests; assessing temporal expectation bias, reliance on language-only shortcuts, video sycophancy, camera motion sensitivity, and robustness to spatiotemporal occlusion. We test leading open- and closed-source VidLMs and find that these models confidently describe reversed scenes as forward, answer questions while neglecting video content, agree with false claims, struggle with basic camera motion, and fail to aggregate temporal information amidst simple spatiotemporal masking. Humans, on the other hand, succeed at these tasks with ease. Alongside our benchmark, we provide a data pipeline that automatically generates diagnostic examples for our stress tests, enabling broader and more scalable evaluation. We will release our benchmark and code to support future research.

视频理解模型评估鲁棒性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。