构建跨语言视频时序推理数据集,测试模型对动作完成性的理解能力。
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
- 设计四语言多选题数据集,结合视频与事件完成标注。
- 顶尖模型在动作持续与完成判断上表现不佳,无法模拟人类推理。
- 适合关注视频理解、时序推理与多模态模型评估的研究者。
人类对事件的感知依赖于区分完成(完成体、有界)与进行(持续体)动作的能力,这一过程受语言结构和视觉线索共同影响。本文提出 extbf{Perfect Times} 数据集,一个全新的四语言(英语、意大利语、俄语、日语)多选题问答基准,用于评估视频-语言模型(VLMs)在时序推理方面的能力。通过将日常活动视频与事件完成标签配对,并设置针对完成体特性的干扰项,该数据集检验模型是否真正理解时序动态,还是仅依赖表面特征。实验结果表明,尽管当前领先模型在文本任务中表现优异,但在基于视频的时序与因果推理方面仍难以达到人类水平。研究强调了整合深层多模态线索以捕捉动作持续与完成细节的重要性,为评估和提升 VLM 的时序推理能力设立了新标准。
原文摘要 · Abstract (English)
Human perception of events is intrinsically tied to distinguishing between completed (perfect and telic) and ongoing (durative) actions, a process mediated by both linguistic structure and visual cues. In this work, we introduce the \textbf{Perfect Times} dataset, a novel, quadrilingual (English, Italian, Russian, and Japanese) multiple-choice question-answering benchmark designed to assess video-language models (VLMs) on temporal reasoning. By pairing everyday activity videos with event completion labels and perfectivity-tailored distractors, our dataset probes whether models truly comprehend temporal dynamics or merely latch onto superficial markers. Experimental results indicate that state-of-the-art models, despite their success on text-based tasks, struggle to mirror human-like temporal and causal reasoning grounded in video. This study underscores the necessity of integrating deep multimodal cues to capture the nuances of action duration and completion within temporal and causal video dynamics, setting a new standard for evaluating and advancing temporal reasoning in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。