arXiv:2503.17899cs.CV2025-03中稿 · TMLR 2025被引 1

从静态图像中学习时间感知,让视觉模型读懂时间线索。

What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images

  • 用跨模态对比学习联合建模图像与时间戳。
  • 在时间估计任务上达到当前最佳性能。
  • 学会的时间感知可应用于图像检索、视频分类等场景。

时间通过光照变化在视觉中显现。受此启发,本文探索从静态图像中学习时间感知的潜力,回答‘时间告诉我们什么’。为此,我们构建了包含130,906张带可靠时间戳的图像的数据集TOC。基于该数据集,提出时间-图像对比学习(TICL)方法,通过跨模态对比学习联合建模时间戳与视觉表征。实验表明,TICL不仅在时间估计任务上超越多个基准指标,且仅通过静态图像训练出的时间感知嵌入,在时间相关下游任务中也表现出色,如基于时间的图像检索、视频场景分类和时间感知图像编辑。结果表明,时间相关的视觉线索可从静态图像中学习,并对多种视觉任务有益,为理解时间相关视觉上下文的研究奠定基础。

原文摘要 · Abstract (English)

Time becomes visible through illumination changes in what we see. Inspired by this, in this paper we explore the potential to learn time awareness from static images, trying to answer: *what time tells us?* To this end, we first introduce a Time-Oriented Collection (TOC) dataset, which contains 130,906 images with reliable timestamps. Leveraging this dataset, we propose a Time-Image Contrastive Learning (TICL) approach to jointly model timestamps and related visual representations through cross-modal contrastive learning. We found that the proposed TICL, 1) not only achieves state-of-the-art performance on the timestamp estimation task, over various benchmark metrics, 2) but also, interestingly, though only seeing static images, the time-aware embeddings learned from TICL show strong capability in several time-aware downstream tasks such as time-based image retrieval, video scene classification, and time-aware image editing. Our findings suggest that time-related visual cues can be learned from static images and are beneficial for various vision tasks, laying a foundation for future research on understanding time-related visual context. Project page: https://rathgrith.github.io/timetells_release/

时间感知图像理解对比学习视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。