arXiv:2411.10019cs.CVcs.LG2024-11被引 2

研究神经网络中间层非极端激活值,发现隐藏的关联信息

Towards Utilising a Range of Neural Activations for Comprehending Representational Associations

  • 分析中等强度神经元激活,而非仅极端值
  • 在合成数据上揭示了最大激活忽略的表征特征
  • 可用于修正真实数据中的虚假关联,适合模型可解释性研究

当前对深度神经网络中间表示的理解多聚焦于通过极端神经元激活和最高方向投影来标注线性方向上的神经元。本文指出,尽管该方法在多数场景下表现良好,却忽略了表示行为中宝贵的信息。神经网络激活通常呈密集分布,更真实的情况是线性方向在不同刺激水平编码信息。我们假设非极端激活包含复杂信息,如统计关联,可能揭示人类可理解的混淆概念。通过研究中间层输出神经元的中等范围激活,我们在合成数据上展示了这些激活能揭示仅靠最大激活无法发现的表征特性。基于此,我们提出一种方法,利用中等范围逻辑值样本重构数据以重训模型,从而缓解真实基准数据集中深层表示中的虚假相关或混淆概念。实验成功验证了考察非最大激活对提取模型学习的复杂关系具有实际价值。

原文摘要 · Abstract (English)

Recent efforts to understand intermediate representations in deep neural networks have commonly attempted to label individual neurons and combinations of neurons that make up linear directions in the latent space by examining extremal neuron activations and the highest direction projections. In this paper, we show that this approach, although yielding a good approximation for many purposes, fails to capture valuable information about the behaviour of a representation. Neural network activations are generally dense, and so a more complex, but realistic scenario is that linear directions encode information at various levels of stimulation. We hypothesise that non-extremal level activations contain complex information worth investigating, such as statistical associations, and thus may be used to locate confounding human interpretable concepts. We explore the value of studying a range of neuron activations by taking the case of mid-level output neuron activations and demonstrate on a synthetic dataset how they can inform us about aspects of representations in the penultimate layer not evident through analysing maximal activations alone. We use our findings to develop a method to curate data from mid-range logit samples for retraining to mitigate spurious correlations, or confounding concepts in the penultimate layer, on real benchmark datasets. The success of our method exemplifies the utility of inspecting non-maximal activations to extract complex relationships learned by models.

神经网络表征分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。