arXiv:2505.14321cs.CV2025-05被引 20

拆解视频大模型评测陷阱,区分真时序理解与伪能力

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

  • 按问题类型自动分类:语言可答、语义可答、时序依赖三类
  • 发现模型在时序问题上表现差,但整体分数虚高
  • 适合做视频理解评估的科研人员和评测设计者

现有视频理解评测常混淆知识性问题与纯图像问题,未能清晰区分模型对动态内容的时序推理能力。我们指出两大问题:一是强语言先验使模型无需观看视频即可作答;二是帧序打乱后模型表现不变,显示其对时间顺序不敏感。为此提出VBenchComp自动化流程,将问题分为四类:LLM-Answerable(无需看视频)、Semantic(帧序打乱仍可答)、Temporal(需正确时序)、Others。该框架实现对视频大模型能力的细粒度评估。分析揭示传统总分掩盖了模型在时序理解上的真实短板,为未来评测设计提供改进方向。

原文摘要 · Abstract (English)

Existing video understanding benchmarks often conflate knowledge-based and purely image-based questions, rather than clearly isolating a model's temporal reasoning ability, which is the key aspect that distinguishes video understanding from other modalities. We identify two major limitations that obscure whether higher scores truly indicate stronger understanding of the dynamic content in videos: (1) strong language priors, where models can answer questions without watching the video; and (2) shuffling invariance, where models maintain similar performance on certain questions even when video frames are temporally shuffled. To alleviate these issues, we propose VBenchComp, an automated pipeline that categorizes questions into different domains: LLM-Answerable, Semantic, and Temporal. Specifically, LLM-Answerable questions can be answered without viewing the video; Semantic questions remain answerable even when the video frames are shuffled; and Temporal questions require understanding the correct temporal order of frames. The rest of the questions are labeled as Others. This can enable fine-grained evaluation of different capabilities of a video LLM. Our analysis reveals nuanced model weaknesses that are hidden by traditional overall scores, and we offer insights and recommendations for designing future benchmarks that more accurately assess video LLMs.

视频理解时序推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。