利用DC分解揭示神经网络内部单调性,提升可解释性
Hidden Monotonicity: Explaining Deep Neural Networks via their DC Decomposition
- 将ReLU网络分解为两个单调凸部分,克服权重爆炸问题
- SplitCAM与SplitLRP在ImageNet-S上优于现有方法,全指标领先
- 构建差值形式的模型,实现内在自解释能力
研究表明,单调性有助于提升神经网络的可解释性。然而,并非所有函数都能被单调神经网络良好逼近。本文展示了单调性仍可通过两种方式增强可解释性:第一,提出一种改进的已训练ReLU网络的DC分解方法,将其拆分为两个单调凸部分,克服了该过程中的数值不稳定问题;基于此,提出的Saliency方法SplitCAM与SplitLRP在ImageNet-S数据集上对VGG16和ResNet18模型,在全部Quantus可解释性评估类别中均取得当前最优表现。第二,我们证明将模型训练为两个单调神经网络之差的形式,能赋予系统强大的自解释特性。
原文摘要 · Abstract (English)
It has been demonstrated in various contexts that monotonicity leads to better explainability in neural networks. However, not every function can be well approximated by a monotone neural network. We demonstrate that monotonicity can still be used in two ways to boost explainability. First, we use an adaptation of the decomposition of a trained ReLU network into two monotone and convex parts, thereby overcoming numerical obstacles from an inherent blowup of the weights in this procedure. Our proposed saliency methods - SplitCAM and SplitLRP - improve on state of the art results on both VGG16 and Resnet18 networks on ImageNet-S across all Quantus saliency metric categories. Second, we exhibit that training a model as the difference between two monotone neural networks results in a system with strong self-explainability properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。