用可解释方法实现东南亚11国图像地理定位,准确率达85.91%。
GeoSEAN: Explainable Country-Level Image Geolocation for ASEAN Regions

- 结合视觉特征与注意力分析,构建可解释的国家级图像定位流程
- 在11个东盟国家数据集上达85.91%准确率和F1值
- 通过物体检测与注意力重播揭示视觉线索的决策依据,适合需要透明性的应用
图像地理定位旨在仅通过视觉内容推断图像的地理位置。然而,在城市景观、道路环境、建筑风格及自然生态相似的地区,该任务仍具挑战性。现有模型多关注坐标预测或分类性能,缺乏对视觉证据如何支撑判断的深入解析。本研究提出面向11个东盟国家的可解释国家级图像地理定位流程。首先,从GeoGuessr类来源、Google Images及街景图像收集4,850张图片;随后在该数据集上评估三种方法:CLIP零样本分类、LightGBM分类器与MLP分类器,其中MLP表现最佳,测试准确率达85.91%,F1值为85.91%。为增强可解释性,采用CLIP注意力回传、YOLO26物体检测及基于能量的指向游戏(EBPG)重叠度量对MLP预测结果进行事后分析。物体级分析显示,高频出现的物体未必对应最高注意力密度,表明频率与注意力反映场景的不同维度。结果表明,该模型不仅支持高精度区域定位,还可实现对视觉线索的物体级审视。
原文摘要 · Abstract (English)
Image geolocation aims to infer the geographic origin of an image from visual content alone. However, this task remains challenging in regions where countries share similar urban, roadside, architectural, and environmental characteristics. Many existing geolocation models focus on coordinate level prediction or classification performance while providing limited insight into how visual evidence contributes to location predictions. This study presents an explainable country level image geolocation pipeline for 11 ASEAN countries. First, we collected 4,850 images from GeoGuessr style sources, Google Images, and additional street level imagery. We then evaluated three approaches on this dataset: CLIP zero shot classification, a LightGBM classifier, and an MLP classifier. The MLP achieved the best test performance, attaining an accuracy and F1 score of 85.91%. For explainability, predictions generated by the MLP classifier were analyzed post hoc using CLIP attention rollout, YOLO26 object detection on the original images, and Energy Based Pointing Game (EBPG) overlap metrics. Object level analysis indicates that frequently detected objects are not necessarily associated with the highest attention density, suggesting that object frequency and attention based visual evidence capture different aspects of a scene. These results demonstrate that the proposed model can support accurate regional image geolocation while enabling object level inspection of the visual cues underlying its predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。