构建真实世界地理时间理解基准,测试模型从图像推断时空信息的能力
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
- 基于1455张全球地面图像,要求模型直接从视觉推断时间与空间属性
- 现有模型在时间推理上表现差,即使微调后仍不理想
- 适合研究多模态推理、地理感知或具身智能的学者使用
地理时间理解是指仅凭视觉输入推断位置、时间及上下文属性的能力,对灾害管理、交通规划、自主导航等应用至关重要。尽管近期视觉语言模型(VLMs)在利用地标和路牌进行图像地理定位方面取得进展,但其对时间信号和物理空间线索的推理能力仍有限。为此,我们提出TimeSpot,一个评估VLM在真实世界中地理时间推理能力的基准。TimeSpot包含来自80个国家的1,455张地面图像,要求结构化预测时间属性(季节、月份、一天中的时段、日照状态)和地理属性(大洲、国家、气候区、环境类型、经纬度),并包含检验现实世界不确定性下物理合理性的时空推理任务。对先进开源与闭源VLM的评估显示性能偏低,尤其在时间推理方面。虽经监督微调有所提升,结果仍不足,凸显实现鲁棒、物理一致的地理时间理解亟需新方法。TimeSpot已公开:https://TimeSpot-GT.github.io。
原文摘要 · Abstract (English)
Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography education. Although recent vision-language models (VLMs) have advanced image geo-localization using cues like landmarks and road signs, their ability to reason about temporal signals and physically grounded spatial cues remains limited. To address this gap, we introduce TimeSpot, a benchmark for evaluating real-world geo-temporal reasoning in VLMs. TimeSpot comprises 1,455 ground-level images from 80 countries and requires structured prediction of temporal attributes (season, month, time of day, daylight phase) and geographic attributes (continent, country, climate zone, environment type, latitude-longitude) directly from visual evidence. It also includes spatial-temporal reasoning tasks that test physical plausibility under real-world uncertainty. Evaluations of state-of-the-art open- and closed-source VLMs show low performance, particularly for temporal inference. While supervised fine-tuning yields improvements, results remain insufficient, highlighting the need for new methods to achieve robust, physically grounded geo-temporal understanding TimeSpot is available at: https://TimeSpot-GT.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。