评测大模型在地理与时间定位上的表现,发现加时间信息能显著提升准确率。
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
- 构建包含120年跨度的全球图像基准,联合评估地点与时间预测。
- 实测8款私有+7款开源大模型,普遍缺乏世界知识和时空推理能力。
- 加入时间线索后定位准确率明显提升,适合研究多模态推理的学者。
地理定位旨在通过图像中的视觉线索推断其拍摄位置。传统方法依赖大规模图像训练取得优异效果。随着多模态大语言模型(MLLM)的兴起,近期研究探索其在地理定位中的应用,得益于更高的精度与可解释性。然而现有基准大多忽略图像固有的时间信息,而时间可进一步约束位置判断。为此,我们提出GTPred,一个新型的地理-时间联合预测基准。该基准包含370张分布于全球、跨越超过120年的图像。我们通过联合评估年份与分层地理位置匹配来衡量MLLM预测性能,并通过精心标注的真实推理链评估中间推理过程。在8款专有和7款开源的MLLM上进行实验表明,尽管具备强大的视觉感知能力,当前模型在世界知识和时空推理方面仍存在明显局限。结果还显示,引入时间信息可显著提升定位性能。
原文摘要 · Abstract (English)
Geo-localization aims to infer the geographic location where an image was captured using observable visual evidence. Traditional methods achieve impressive results through large-scale training on massive image corpora. With the emergence of multi-modal large language models (MLLMs), recent studies have explored their applications in geo-localization, benefiting from improved accuracy and interpretability. However, existing benchmarks largely ignore the temporal information inherent in images, which can further constrain the location. To bridge this gap, we introduce GTPred, a novel benchmark for geo-temporal prediction. GTPred comprises 370 globally distributed images spanning over 120 years. We evaluate MLLM predictions by jointly considering year and hierarchical location sequence matching, and further assess intermediate reasoning chains using meticulously annotated ground-truth reasoning processes. Experiments on 8 proprietary and 7 open-source MLLMs show that, despite strong visual perception, current models remain limited in world knowledge and geo-temporal reasoning. Results also demonstrate that incorporating temporal information significantly enhances location inference performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。