用隐式神经表示生成可解释的视觉归因图,让模型决策更透明。
Generating visual explanations from deep networks using implicit neural representations
- 用坐标隐式网络重构极端扰动法,生成符合区域约束的归因掩码。
- 提出迭代方法,可生成同一图像上多个不重叠的归因区域。
- 揭示模型不仅关注物体本身,还依赖常见纹理与背景特征。
让深度学习模型的决策过程对人类可理解,是负责任AI的关键。归因方法是可解释性研究的重要方向,旨在识别输入中影响模型输出最关键的区域。本文表明,隐式神经表示(INRs)是生成视觉归因图的理想框架。首先,我们利用基于坐标的隐式网络重构并扩展了极端扰动技术,生成归因掩码;实验显示,通过合理调节隐式网络,可生成满足指定面积约束的优质掩码。其次,我们提出一种基于INR的迭代方法,能够为同一图像生成多个互不重叠的归因掩码。结果表明,深度学习模型在判定图像标签时,不仅关注目标物体的外观,也依赖于常伴随该物体的区域和纹理。研究表明,隐式网络非常适合生成归因掩码,并能揭示深度学习模型性能背后的内在机制。
原文摘要 · Abstract (English)
Explaining deep learning models in a way that humans can easily understand is essential for responsible artificial intelligence applications. Attribution methods constitute an important area of explainable deep learning. The attribution problem involves finding parts of the network's input that are the most responsible for the model's output. In this work, we demonstrate that implicit neural representations (INRs) constitute a good framework for generating visual explanations. Firstly, we utilize coordinate-based implicit networks to reformulate and extend the extremal perturbations technique and generate attribution masks. Experimental results confirm the usefulness of our method. For instance, by proper conditioning of the implicit network, we obtain attribution masks that are well-behaved with respect to the imposed area constraints. Secondly, we present an iterative INR-based method that can be used to generate multiple non-overlapping attribution masks for the same image. We depict that a deep learning model may associate the image label with both the appearance of the object of interest as well as with areas and textures usually accompanying the object. Our study demonstrates that implicit networks are well-suited for the generation of attribution masks and can provide interesting insights about the performance of deep learning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。