arXiv:2507.05916cs.CVcs.AI2025-07中稿 · IEEE Journal of Se…被引 6

针对遥感图像分类,评估了10种可解释AI方法与指标的有效性。

On the Effectiveness of Methods and Metrics for Explainable AI in Remote Sensing Image Scene Classification

  • 对比5类指标与5种方法在3个遥感数据集上的表现
  • 发现梯度法在多标签图像中失效,基线选择影响扰动法结果
  • 推荐使用鲁棒性与随机化指标,避免误判大范围地物

可解释人工智能(xAI)在遥感(RS)图像场景分类中的发展备受关注。当前多数xAI方法及评价指标源自自然图像的计算机视觉领域,直接应用于遥感图像可能不适用。本文系统分析了五类解释指标(忠实性、鲁棒性、定位性、复杂性、随机化)和五种特征归因方法(Occlusion、LIME、GradCAM、LRP、DeepLIFT)在三个遥感数据集上的有效性。方法论分析揭示:基于扰动的方法受基线和空间特征影响显著;梯度法在多标签图像中表现不佳;部分相关性传播方法(LRP)会不均衡分配相关性。评价指标方面,忠实性与定位性指标对大空间范围类别不可靠;而鲁棒性与随机化指标更稳定。实验结果验证了上述发现,并提供了适用于遥感场景分类的解释方法与指标选择指南。

原文摘要 · Abstract (English)

The development of explainable artificial intelligence (xAI) methods for scene classification problems has attracted great attention in remote sensing (RS). Most xAI methods and the related evaluation metrics in RS are initially developed for natural images considered in computer vision (CV), and their direct usage in RS may not be suitable. To address this issue, in this paper, we investigate the effectiveness of explanation methods and metrics in the context of RS image scene classification. In detail, we methodologically and experimentally analyze ten explanation metrics spanning five categories (faithfulness, robustness, localization, complexity, randomization), applied to five established feature attribution methods (Occlusion, LIME, GradCAM, LRP, and DeepLIFT) across three RS datasets. Our methodological analysis identifies key limitations in both explanation methods and metrics. The performance of perturbation-based methods, such as Occlusion and LIME, heavily depends on perturbation baselines and spatial characteristics of RS scenes. Gradient-based approaches like GradCAM struggle when multiple labels are present in the same image, while some relevance propagation methods (LRP) can distribute relevance disproportionately relative to the spatial extent of classes. Analogously, we find limitations in evaluation metrics. Faithfulness metrics share the same problems as perturbation-based methods. Localization metrics and complexity metrics are unreliable for classes with a large spatial extent. In contrast, robustness metrics and randomization metrics consistently exhibit greater stability. Our experimental results support these methodological findings. Based on our analysis, we provide guidelines for selecting explanation methods, metrics, and hyperparameters in the context of RS image scene classification.

可解释AI遥感图像特征归因评价指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。