arXiv:2603.24749cs.CV2026-03中稿 · CVPR被引 2

统一建模时间、图像与位置,实现按时间和地点精准检索图片。

TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval

  • 构建统一框架,支持单模态或多模态查询,联合学习时空特征。
  • 在时间预测和时空检索任务上分别提升16%、8%和14%的准确率。
  • 适用于数字取证、城市监测等需时空联合推理的场景。

数字取证、城市监测与环境分析等现实应用需同时理解视觉外观、地理位置与时间信息。除传统地理定位与拍摄时间预测外,这些应用日益需要更复杂的检索能力,如根据查询图像在指定时间的相同地点找回对应图像。本文提出TIGeR统一框架,解决「时空感知图像检索」问题,支持单模态与多模态输入,使用同一表征完成(1)地理定位、(2)拍摄时间预测、(3)时空感知图像检索。TIGeR通过保留场景在外观变化下的位置身份,实现基于“何时何地”而非仅“视觉相似”的检索。为此,设计多阶段数据清洗流程,并构建包含450万对图像-位置-时间三元组的训练集及8.6万条高质量评估三元组的数据集。大量实验表明,TIGeR在年周期、日周期预测以及时空检索召回率上均显著优于基线方法,最高提升达16%、8%与14%,验证了统一时空建模的有效性。

原文摘要 · Abstract (English)

Many real-world applications in digital forensics, urban monitoring, and environmental analysis require jointly reasoning about visual appearance, location, and time. Beyond standard geo-localization and time-of-capture prediction, these applications increasingly demand more complex capabilities, such as retrieving an image captured at the same location as a query image but at a specified target time. We formalize this problem as Geo-Time Aware Image Retrieval and propose TIGeR, a unified framework for Time, Images and Geo-location Retrieval. TIGeR supports flexible input configurations (single-modality and multi-modality queries) and uses the same representation to perform (i) geo-localization, (ii) time-of-capture prediction, and (iii) geo-time-aware retrieval. By preserving the underlying location identity despite large appearance changes, TIGeR enables retrieval based on where and when a scene was captured, rather than purely on visual similarity. To support this task, we design a multistage data curation pipeline and propose a new diverse dataset of 4.5M paired image-location-time triplets for training and 86k high-quality triplets for evaluation. Extensive experiments show that TIGeR consistently outperforms strong baselines and state-of-the-art methods by up to 16% on time-of-year, 8% time-of-day prediction, and 14% in geo-time aware retrieval recall, highlighting the benefits of unified geo-temporal modeling.

图像检索时空建模多模态数字取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。