arXiv:2605.00912cs.CV2026-05

通过可解释性分析,揭示图像定位模型依赖的具体物体线索。

Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case

论文配图:Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case
图 1 · 摘自论文原文
  • 从注意力图提取显著区域并分割为类物体元素
  • 引导裁剪区域比随机区域保留更多预测信息(三国家基准)
  • 帮助理解模型是否像人一样依赖具体视觉线索

人类在玩地理定位游戏(如GeoGuessr)时,会依赖道路标线、植被或建筑细节等具体视觉线索判断拍摄位置。现有图像地理定位模型是否也使用类似物体级证据尚不明确,因为梯度类方法(如Grad-CAM)通常只生成模糊区域,难以关联到具体物体或可感知模式。本文提出一种以物体为中心的分析流程:从归因图中提取显著区域并分割为类物体元素,通过删除与插入测试评估其预测相关性,对比归因引导裁剪与覆盖相似的随机区域。在三国家基准上的实验表明,归因引导裁剪始终比随机裁剪保留更多模型预测信息。结果表明,归因图可被分解为可解释、可感知的视觉元素,为地理定位模型的物体级分析提供新路径。

原文摘要 · Abstract (English)

When humans play geolocation games such as GeoGuessr, they rely on concrete visual cues, such as road markings, vegetation, or architectural details, to infer where an image was captured. Whether image geolocation models rely on similar object-level evidence remains difficult to determine, as attribution methods like Grad-CAM typically highlight diffuse regions rather than coherent visual entities, making it difficult to link model predictions to specific objects or perceptible patterns. In this work, we propose an object-centric analysis pipeline to investigate the visual evidence used by geolocation models. Starting from attribution maps, we extract salient regions and segment them into object-like elements. We evaluate their predictive relevance through deletion and insertion tests, comparing attributionguided crops to randomly selected regions with similar coverage. Experiments on a three-country benchmark show that attribution-guided crops consistently retain more information for the model's prediction than random crops. These results suggest that attribution maps can be decomposed into interpretable, perceptible elements, providing a step toward object-level analysis of geolocation models.

可解释性图像定位物体级分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。