通过多层特征图融合,提升图像地理定位的可解释性。
Combi-CAM: A Novel Multi-Layer Approach for Explainable Image Geolocalization

- 结合网络多层梯度加权特征图,而非仅用深层特征。
- 实现对图像不同区域贡献度的精细化分析。
- 适合需要理解模型决策过程的研究者与应用者。
全球尺度的图像地理定位任务旨在仅基于图像视觉特征估算其拍摄位置。尽管深度学习模型(尤其是卷积神经网络,CNN)在此领域取得显著进展,但理解其预测背后的推理过程仍具挑战性。本文提出Combi-CAM,一种新型多层可解释方法,通过融合网络多个层级的梯度加权类激活图(Grad-CAM),而非传统仅依赖最深层特征的方式,使模型决策过程的可视化更精细。该方法能够揭示图像中不同视觉特征对最终定位结果的贡献程度,提供比现有方法更深入的解释能力。
原文摘要 · Abstract (English)
Planet-scale photo geolocalization involves the intricate task of estimating the geographic location depicted in an image purely based on its visual features. While deep learning models, particularly convolutional neural networks (CNNs), have significantly advanced this field, understanding the reasoning behind their predictions remains challenging. In this paper, we present Combi-CAM, a novel method that enhances the explainability of CNN-based geolocalization models by combining gradient-weighted class activation maps obtained from several layers of the network architecture, rather than using only information from the deepest layer as is typically done. This approach provides a more detailed understanding of how different image features contribute to the model's decisions, offering deeper insights than the traditional approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。