arXiv:2510.26241cs.CVcs.CL2025-10被引 1

测试视觉语言模型判断视频时间方向的能力,发现多数表现接近随机。

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

  • 用人类心理物理实验验证的基准测试模型对时间方向的判断能力。
  • 最佳模型在不可逆物理过程上准确率远低于人类,如自由落体、爆炸扩散。
  • 适合关注视频时序理解、物理推理的AI研究者使用。

现代视觉语言模型(VLM)在多模态任务中表现优异,但对视频中的时间信息理解仍较薄弱,且缺乏有效评估。我们通过一个看似简单却极具揭示性的问题来填补这一空白:判断一段短视频是正放还是倒放(时间箭头,AoT)。为此,我们构建了AoT-PsyPhyBENCH,一个经过心理物理学验证的基准,使用与人类实验相同的刺激和行为基线,评估VLM在自然视频中推断时间方向的能力。对开源与专有、推理型与非推理型VLM的全面评估表明,大多数模型表现接近随机,即使最优模型在物理不可逆过程(如自由落体、扩散/爆炸)和因果性手动动作(如分割/相加)上,准确率也远低于人类,而这些现象人类可瞬间识别。结果揭示当前多模态系统的核心缺陷:虽能捕捉丰富的视觉-语义关联,却缺乏时间连续性和因果理解所需的归纳偏置。我们公开发布AoT-PsyPhyBENCH的代码与数据,以推动VLM在物理与时间推理方面的发展。

原文摘要 · Abstract (English)

Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but revealing challenge: judging the arrow of time (AoT)-whether a short clip is played forward or backward. We introduce AoT-PsyPhyBENCH, a psychophysically validated benchmark that tests whether VLMs can infer temporal direction in natural videos using the same stimuli and behavioral baselines established for humans. Our comprehensive evaluation of open-weight and proprietary, reasoning and non-reasoning VLMs reveals that most models perform near chance, and even the best model lags far behind human accuracy on physically irreversible processes (e.g., free fall, diffusion/explosion) and causal manual actions (division/addition) that humans recognize almost instantly. These results highlight a fundamental gap in current multimodal systems: while they capture rich visual-semantic correlations, they lack the inductive biases required for temporal continuity and causal understanding. We release the code and data for AoT-PsyPhyBENCH to encourage further progress in the physical and temporal reasoning capabilities of VLMs.

时间推理视觉语言模型物理理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。