评估单目深度估计模型的可解释性,找出关键输入区域。
Shedding Light on Depth: Explainability Assessment in Monocular Depth Estimation
- 用显著图和积分梯度分析深度模型的关键输入特征。
- 轻量级与深度模型中,显著图与积分梯度表现良好。
- 提出新指标Attribution Fidelity,更可靠识别无效解释。
可解释人工智能在理解深度学习模型决策过程、提升可信度方面日益重要。然而,尽管单目深度估计(MDE)在实际应用中广泛部署,其可解释性仍缺乏研究。本文研究如何分析MDE网络,将输入图像映射到预测深度图。具体考察了显著图、积分梯度和注意力传播等经典特征归因方法,在METER(轻量级)与PixelFormer(深度)模型上的表现。通过选择性扰动归因方法识别的重要与不重要像素,分析对模型输出的影响来评估解释质量。由于现有评价指标在衡量MDE视觉解释有效性时存在局限,本文引入新的“归因保真度”(Attribution Fidelity)指标,评估归因一致性与预测深度图的匹配程度。实验表明,显著图在轻量级模型中表现优异,积分梯度在深度模型中更优;且该新指标能有效识别传统指标误判为有效的不可靠解释。
原文摘要 · Abstract (English)
Explainable artificial intelligence is increasingly employed to understand the decision-making process of deep learning models and create trustworthiness in their adoption. However, the explainability of Monocular Depth Estimation (MDE) remains largely unexplored despite its wide deployment in real-world applications. In this work, we study how to analyze MDE networks to map the input image to the predicted depth map. More in detail, we investigate well-established feature attribution methods, Saliency Maps, Integrated Gradients, and Attention Rollout on different computationally complex models for MDE: METER, a lightweight network, and PixelFormer, a deep network. We assess the quality of the generated visual explanations by selectively perturbing the most relevant and irrelevant pixels, as identified by the explainability methods, and analyzing the impact of these perturbations on the model's output. Moreover, since existing evaluation metrics can have some limitations in measuring the validity of visual explanations for MDE, we additionally introduce the Attribution Fidelity. This metric evaluates the reliability of the feature attribution by assessing their consistency with the predicted depth map. Experimental results demonstrate that Saliency Maps and Integrated Gradients have good performance in highlighting the most important input features for MDE lightweight and deep models, respectively. Furthermore, we show that Attribution Fidelity effectively identifies whether an explainability method fails to produce reliable visual maps, even in scenarios where conventional metrics might suggest satisfactory results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。