arXiv:2409.08712cs.LGcs.AI2024-09ICML被引 7

首次量化神经网络各层新出现与被遗忘的交互模式。

Layerwise Change of Knowledge in Neural Networks

  • 扩展交互定义,提取中间层编码的交互模式。
  • 发现深层网络逐步增强新交互、抑制噪声交互。
  • 揭示模型泛化能力与特征不稳定性随层变化规律。

本文旨在解释深度神经网络(DNN)在前向传播过程中如何逐层提取新知识并遗忘噪声特征。尽管目前对DNN所编码知识的定义尚未达成共识,但已有研究通过数学证据表明,交互(interactions)是DNN编码的符号化推理模式的基本单元。本文扩展了交互的定义,并首次从中间层中提取出交互模式。通过量化并追踪每一层中新出现的交互和被遗忘的交互,揭示了DNN学习行为的新机制。层间交互变化还反映了模型泛化能力与特征表示不稳定性随网络深度的变化过程。

原文摘要 · Abstract (English)

This paper aims to explain how a deep neural network (DNN) gradually extracts new knowledge and forgets noisy features through layers in forward propagation. Up to now, although the definition of knowledge encoded by the DNN has not reached a consensus, Previous studies have derived a series of mathematical evidence to take interactions as symbolic primitive inference patterns encoded by a DNN. We extend the definition of interactions and, for the first time, extract interactions encoded by intermediate layers. We quantify and track the newly emerged interactions and the forgotten interactions in each layer during the forward propagation, which shed new light on the learning behavior of DNNs. The layer-wise change of interactions also reveals the change of the generalization capacity and instability of feature representations of a DNN.

神经网络知识提取交互分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。