arXiv:2410.02763cs.CVcs.AI2024-10被引 34

测试发现大模型在短视频时间推理上仍严重不足

Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos

  • 构建1000对自然短视频-文本对,评估时间因果推理能力
  • 最佳模型GPT-4o仅达50%准确率,远低于人类90%基准
  • 开源数据集与代码,适合评估视频理解模型

近期普遍认为现代大型多模态模型已解决短视频理解的核心挑战,学术界与产业界正转向长视频理解的复杂问题。然而事实是否如此?我们的研究显示,即使面对短视频,大模型仍缺乏基本的时间推理能力。为此,我们提出了Vinoground——一个包含1000对自然短视频与对应文本的时序反事实评估基准。实验表明,现有大模型难以区分不同动作及物体变化的时间差异。例如,最优模型GPT-4o在文本与视频评分上仅达到约50%,远低于人类基准的约90%。所有开源多模态模型与CLIP基模型表现更差,接近随机猜测水平。本工作揭示:短视频时间推理仍是未被解决的关键难题。数据集与评估代码已公开于https://vinoground.github.io。

原文摘要 · Abstract (English)

There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their attention towards the more complex challenges posed by understanding long-form videos. However, is this really the case? Our studies indicate that LMMs still lack many fundamental reasoning capabilities even when dealing with short videos. We introduce Vinoground, a temporal counterfactual LMM evaluation benchmark encompassing 1000 short and natural video-caption pairs. We demonstrate that existing LMMs severely struggle to distinguish temporal differences between different actions and object transformations. For example, the best model GPT-4o only obtains ~50% on our text and video scores, showing a large gap compared to the human baseline of ~90%. All open-source multimodal models and CLIP-based models perform much worse, producing mostly random chance performance. Through this work, we shed light onto the fact that temporal reasoning in short videos is a problem yet to be fully solved. The dataset and evaluation code are available at https://vinoground.github.io.

视频理解时间推理多模态模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。