测试视觉语言模型对视频时序一致性的理解能力,发现其严重不足。
What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

- 将时序异常建模为异常检测任务,通过帧交换和噪声替换构建测试数据。
- 模型在单帧异常检测中表现良好,但在时序异常检测上接近随机水平。
- 人类表现远超模型,提示当前模型缺乏跨帧推理能力,适合研究时序理解的学者。
视觉语言模型(VLMs)在视频和图像序列基准上表现优异,但其是否真正捕捉了时序结构仍不明确。为此,我们提出将时序定位建模为异常检测问题,设计了一个简单且可控的评估框架,直接检验模型对时序一致性的敏感性。引入TimeCatch:通过交换连续帧制造时序异常,用高斯噪声替换帧生成帧级异常。在四个合成与真实数据集上评估模型在异常检测与定位任务的表现,并开展人类对照实验。结果表明,模型在帧级异常检测上表现稳定且定位准确,但在时序异常检测上仅接近随机水平,定位表现也仅为略高于随机。人类在两项任务中均接近满分。额外分析显示,模型规模、提示策略、序列长度及视觉相似性等因素无法完全解释失败原因。这些发现表明,当前VLMs能识别单帧异常,却难以整合多帧信息进行时序推理。TimeCatch为评估视觉语言模型的时序理解能力提供了受控基准。
原文摘要 · Abstract (English)
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection. While VLMs consistently detect frame-level anomalies and often localize them accurately, they perform near chance on temporal anomaly detection and only modestly above chance on localization. Humans, in contrast, achieve near-ceiling performance on both tasks. Additional analyses across model scales, prompting strategies, sequence lengths, and visual similarity suggest that these failures cannot be explained solely by limitations in perception or model capacity. Together, these findings indicate that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency. TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。