用超像素分组提升神经网络解释图的稳定性和可读性
A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
- 通过超像素对像素分组,降低梯度类解释图的随机波动
- 在CIFAR-10和ImageNet上,解释结果稳定性显著优于传统像素级方法
- 适合需要可重复、可信赖模型解释的视觉领域研究者
显著性图在计算机视觉中被广泛用于解释神经网络分类器。然而,由于训练样本和优化算法的随机性,生成的显著性图存在较高程度的随机性,导致领域专家难以捕捉影响模型决策的内在因素。本文提出一种新颖的像素分组策略,以提升基于梯度的显著性图的稳定性和泛化能力。通过理论分析与数值实验,我们证明像素分组能有效降低显著性图的方差,并改善解释方法的泛化行为。进一步地,我们提出基于超像素的合理分组策略,将像素聚类为与图像语义一致的区域。我们在CIFAR-10和ImageNet数据集上进行了多组实验,结果表明,基于超像素的解释图在稳定性和质量上均持续优于基于像素的显著性图。
原文摘要 · Abstract (English)
Saliency maps are widely used in the computer vision community for interpreting neural network classifiers. However, due to the randomness of training samples and optimization algorithms, the resulting saliency maps suffer from a significant level of stochasticity, making it difficult for domain experts to capture the intrinsic factors that influence the neural network's decision. In this work, we propose a novel pixel partitioning strategy to boost the stability and generalizability of gradient-based saliency maps. Through both theoretical analysis and numerical experiments, we demonstrate that the grouping of pixels reduces the variance of the saliency map and improves the generalization behavior of the interpretation method. Furthermore, we propose a sensible grouping strategy based on super-pixels which cluster pixels into groups that align well with the semantic meaning of the images. We perform several numerical experiments on CIFAR-10 and ImageNet. Our empirical results suggest that the super-pixel-based interpretation maps consistently improve the stability and quality over the pixel-based saliency maps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。