arXiv:2510.19559cs.CVcs.AI2025-10被引 1

发现视觉语言模型中时间信息有结构,可提取时间线用于推理

A Matter of Time: Revealing the Structure of Time in Vision-Language Models

  • 通过新数据集和方法揭示时间在模型嵌入空间中的非线性低维结构
  • 基于该结构构建显式时间线表示,准确率优于提示基线且高效
  • 适合研究多模态时间推理或模型可解释性的研究人员

大规模视觉语言模型(如CLIP)凭借多样文本元数据获得开放词汇能力,能解决训练范围外的任务。本文研究这类模型对时间的感知能力,引入包含超过10,000张图像与时间真实标签的TIME10k基准数据集,评估37个VLM的时间感知能力。结果表明,时间信息在模型嵌入空间中呈现低维、非线性流形结构。基于此,提出从嵌入空间中推导显式时间线表示的方法,该表示能建模时间及其连续性,从而支持时间推理任务。所提时间线方法在准确性上达到甚至超过提示基线,同时计算开销小。代码与数据已公开于https://tekayanidham.github.io/timeline-page/。

原文摘要 · Abstract (English)

Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ``timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.

视觉语言模型时间推理嵌入结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。