通过模块化训练提升神经网络可解释性
Training Neural Networks for Modularity aids Interpretability
- 设计纠缠损失函数,引导模型形成互不干扰的模块
- 在CIFAR-10上发现各模块学习独立且更小的特征电路
- 适合关注模型内部机制与可解释性的研究者
提升网络可解释性的方法之一是通过可聚类性,即把模型拆分为互不重叠的模块以便独立分析。我们发现预训练模型高度不可聚类,因此使用一种‘纠缠损失’函数训练模型,促进非交互模块的形成。通过自动化可解释性度量,我们证明该方法在CIFAR-10上能识别出学习不同、独立且更小电路的模块。该方法为使神经网络更易解释提供了有前景的方向。
原文摘要 · Abstract (English)
An approach to improve network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We find pretrained models to be highly unclusterable and thus train models to be more modular using an ``enmeshment loss'' function that encourages the formation of non-interacting clusters. Using automated interpretability measures, we show that our method finds clusters that learn different, disjoint, and smaller circuits for CIFAR-10 labels. Our approach provides a promising direction for making neural networks easier to interpret.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。