arXiv:2606.05702cs.AIcs.CV2026-06

测试视觉语言模型对时间顺序的理解能力,发现它们常依赖颜色等表面线索。

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

论文配图:Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models
图 1 · 摘自论文原文
  • 构建三类新数据集,涵盖历史跨度、事件类型和图文时间对齐
  • 模型在多数场景表现不佳,且易受颜色等错误线索干扰
  • 适合研究多模态推理与模型偏差的学者参考

近期视觉语言模型(VLMs)在理解复杂视觉语义方面取得显著进展,但其对时间顺序的推理能力仍待深入探索。本文提出一个全新基准,专门评估VLMs对图像内及跨图像的时间信息感知与推理能力。不同于以往以视频帧序列为焦点的基准,本研究深入分析时间判断背后的逻辑,并拓展至多模态融合。为此,我们构建三个专用数据集:包含跨越长期历史的视觉相似物体、按多样化事件与物体类型分类的数据,以及将图像与时间敏感新闻文本配对的跨模态数据。通过大量实验,分析模型在不同类别中的性能差异,并重点考察其是否依赖‘错误捷径’,如图像色彩而非真实时间特征。结果表明,尽管VLMs展现出潜力,却常利用灰度与彩色滤镜等表面线索绕过真正的时序推理。我们提供高质量数据集与严谨评估框架,为识别当前局限与推动更稳健、逻辑一致的多模态模型发展提供诊断工具。源代码见https://github.com/LuoRenqiang/ChronoVision。

原文摘要 · Abstract (English)

Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored. In this paper, we introduce a novel benchmark specifically designed to evaluate how VLMs perceive and reason about chronological information within and across images. Unlike existing video-based benchmarks that focus on frame sequencing, our work delves into the underlying logic of chronological judgment and the expansion toward multimodal integration. To facilitate this, we construct three specialized datasets: one containing visually similar objects spanning long historical durations, another categorized by diverse event and object types, and a third pairing images with time-sensitive news text for cross-modal alignment. Through extensive experiments, we analyze whether models exhibit performance disparities across categories and, crucially, explore whether they rely on ``incorrect shortcuts'', such as image color rather than genuine chronological features. Our results reveal that while VLMs show promise, they frequently exploit superficial cues like grayscale versus color filters to bypass authentic chronological reasoning. By providing these high-quality datasets and a rigorous evaluation framework, we offer a diagnostic tool to identify current limitations and guide the development of more robust, logically grounded multimodal models. The source code is shown in https://github.com/LuoRenqiang/ChronoVision.

视觉语言模型时间推理模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。